View by category
How are the names of Topics and Topic Clusters generated?
Last updated on November 15, 2024
Also available in
Topic names are produced in a data-driven fashion based on keyphrases extracted from the titles, abstracts and author keywords associated to a Topic. Keyphrases are obtained by annotating content with OmniScience thesaurus.
OmniScience is an in-house developed single unified thesaurus spanning all major disciplines to create a list of standardized concepts. OmniScience thesaurus consists of more than 50,000+ manually curated core concepts and more than 700,000 semi-automated created extension concepts which are more fine-grained.
For Topic names we use two extension and one core concept. For Topic Cluster names we use three core concepts.
The concepts (keyphrases) that constitute a Topic / Topic Cluster name were chosen to be highly frequent to the Topic / Topic Cluster, whilst at the same time being specific enough to the Topic / Topic Cluster to obtain a meaningful and unique name.
We also apply additional rules so that the names contain fewer redundant terms - e.g. excluding “Image Analysis” when the first three most frequent concepts are “Image Processing; Image Analysis; Strain”.
In 2024, papers in the Topics published as of 2013 were used in the naming exercise. If there were less than 50 publications in that time range, the 20-year range was used instead.
Did we answer your question?
Related answers
Recently viewed answers
Functionality disabled due to your cookie preferences
Back to Topics FAQs