What does "Relevance" mean in Scopus?
Last updated on August 15, 2024
In Scopus, it is possible to sort a list of documents by a date, author, source title, citedness, or by "relevance".
By default, the results are sorted on date, with the most recently published first.
"Relevance" is a term that often causes confusion. If you or I say "I found some really relevant papers on Scopus", it means that we found articles that match the idea we had in our minds when we started searching. Relevance for humans is something subjective.
What is search engine relevance?
Search engine relevance, on the other hand, is simply a statistical calculation, albeit a very intricate and sophisticated one. This calculation tells us how well text in the documents returned as search results reflect the terms and criteria executed in a search query.
How is search engine relevance calculated on Scopus?
Our search engine uses a sophisticated relevance model based on proven concepts from the science of Information Retrieval and long experience with Web, data, and enterprise searching. Each document returned as a result by a query is given a score composed of a variety of different factors. The result of the calculation is dependent upon the query itself, on the documents being searched, and on the way the these two different factors are balanced against one another.
Sorting on this rank score helps us evaluate and select within large sets of documents. In some searches, the articles first in the list might indeed be the most relevant to the topic we have in mind; in others, seeing what rises to the top may give us insights into how we can alter and refine our search to get better results.
The following factors contribute to the rank score as part of the calculation:
Factor |
Calculation |
|---|---|
Number of hits | The more often a term occurs in a document, the more likely that it is relevant to the topic of the article. So the number of times a term appears in a document plays a role in calculating relevancy and is described with more detail in the sections below. |
How significant is a word? | Not every word is equally important. A term that occurs in nearly all documents will score less than something unusual. We use calculations based on Term Frequency/Inverse. Scopus uses Document Frequency (TF/IDF) (a concept originally introduced by Karen Spärck Jones, 1972) to assign a weight to any particular word in any collection of documents - in this case all the documents in Scopus. For example: The word "experiment" is likely to be found quite commonly in all subject areas in Scopus, whereas a word like "crocodile" or "Cryptococcus" is less common across the whole range of subjects and will have more impact on ranking. |
Where is the word found? | It is also important where terms occur in the document. If a word is in the title, abstract, or keywords in a scientific article, then it is probably very important. If it is only mentioned in the references, it probably is not very important. To reflect this contrast, Scopus assigns different weights to hits in different sections of the document. |
Position in the document | The earlier in the text a term occurs, the more important it is likely to be. The position of the first occurrence therefore makes a contribution to the ranking. |
Proximity | The closer together different terms from the query are found, the higher the score. |
Completeness | If a query consists of several words, then a document that contains all the words within the same field will score higher than one in which the words are scattered throughout the document. |
With the detailed searching that is possible in Scopus, it is not always possible to apply all the relevant calculations listed above. Notably: Terms with a wildcard cannot contribute to the ranking because it is not a single term, but a set of possibilities. Remember, going back to our definition above, that some results lists on Scopus cannot be ranked by search engine relevance because they do not come from a text search. An example of this would be a list of articles that cite another article.
We have tuned the relative contributions of the factors described above to give an optimal mix across a wide range of searches and topics. The relevance model does allow for more contributing factors than Scopus presently employs. For example, citedness, or popularity, or date. However, adding these non-textual measures to the calculation would require us to make judgments or assumptions on entered search criteria. For this reason they are not included.
References
Spärck Jones, Karen (1972) "A statistical interpretation of term specificity and its application in retrieval". Journal of Documentation 28 (1): 11-21. doi:10.1108/eb026526.
Did we answer your question?
Related answers
Recently viewed answers
Functionality disabled due to your cookie preferences