How are the names of Topics and Topic Clusters generated?

Last updated on November 15, 2024

Also available in

Topic names are produced in a data-driven fashion based on keyphrases extracted from the titles, abstracts and author keywords associated to a Topic. Keyphrases are obtained by annotating content with OmniScience thesaurus.

OmniScience is an in-house developed single unified thesaurus spanning all major disciplines to create a list of standardized concepts. OmniScience thesaurus consists of more than 50,000+ manually curated core concepts and more than 700,000 semi-automated created extension concepts which are more fine-grained.

For Topic names we use two extension and one core concept. For Topic Cluster names we use three core concepts.

The concepts (keyphrases) that constitute a Topic / Topic Cluster name were chosen to be highly frequent to the Topic / Topic Cluster, whilst at the same time being specific enough to the Topic / Topic Cluster to obtain a meaningful and unique name.

We also apply additional rules so that the names contain fewer redundant terms - e.g. excluding “Image Analysis” when the first three most frequent concepts are “Image Processing; Image Analysis; Strain”.

In 2024, papers in the Topics published as of 2013 were used in the naming exercise. If there were less than 50 publications in that time range, the 20-year range was used instead.

Did we answer your question?

Related answers

Recently viewed answers

Functionality disabled due to your cookie preferences

Also available in

For further assistance: