View by category
Top-Cited Scientists Database: FAQs
Last updated on October 07, 2026
This FAQ answers common questions about the Top-Cited Scientists database, including how it is created, how to understand inclusion and citation metrics, and how to request corrections to Scopus author profiles. It also explains where to get help with the database and related support services. Each edition is based on a fixed Scopus snapshot and cannot be changed after publication.
FREQUENTLY ASKED QUESTIONS
Please check carefully the answers to these frequently asked questions and the provided references. We will not answer any e-mails with questions that are covered by material presented in the Frequently Asked Questions or in the introductory documentation to the databases.
This is a publicly available database of highly cited scientists, built from Scopus data. It provides standardized citation metrics for each scientist, including total citations, h-index, a co-authorship-adjusted hm-index, citations broken down by authorship position (single, first, last author), and a composite indicator called the c-score. Separate figures are given for career-long impact and for a single recent calendar year, and metrics are shown both with and without self-citations. The dataset also flags retracted papers (based on the Retraction Watch database) and citations to and from retracted papers. Scientists are classified into 20 scientific fields and 174 sub-fields using the Science-Metrix classification, and field- and subfield-specific percentiles are provided for every scientist with at least 5 papers. Inclusion is based on ranking in the top 100,000 scientists by c-score (with or without self-citations), or a percentile rank of 2% or higher within a sub-field.
The dataset is calculated from Scopus author profiles as they exist on a fixed snapshot date. For example, version 9 is based on the August 1, 2026 Scopus snapshot, with career-long metrics updated through end-of-2025 and single-year metrics reflecting citations received during calendar year 2025. The c-score, the composite indicator used to rank scientists, focuses on citation impact rather than raw publication counts, and factors in co-authorship patterns and author position (single, first, or last author). A new version of the dataset is produced annually, each based on a fresh Scopus snapshot.
Being absent from the list simply means the composite indicator (c-score) value was not high enough to meet the inclusion threshold for that version, either the top 100,000 scientists overall or the top 2% within a sub-field. It does not mean the author's work is not of good quality.
Why has my ranking and/or inclusion status changed?
Each version of the dataset reflects a new Scopus snapshot and an additional year of citation data, so movement in rank is expected. Typical causes include new citations accrued by the scientist and by others in the field, changes to the scientist's own Scopus profile (papers added, merged, or reassigned), shifts in field/sub-field classification and updates to self-citation calculations. Because rankings are relative to all other scientists in the database, even an increase in one's own citation counts can coincide with a lower or higher rank depending on how the wider field has moved.
Can the dataset be corrected after publication?
No. Each version of the dataset is published in archival form and is not changed after release. If a correction is submitted, even if it reaches the correct team, it will only be reflected in the next annual version of the dataset, not in the version already published.
How can I improve the accuracy of my Scopus profile?
Since the dataset reflects Scopus author profiles as they stood on the snapshot date, keeping that profile accurate and complete is the best way to ensure it is represented correctly in future versions. Requests for corrections to Scopus data, including author names, publication lists, and affiliations, should not be sent to the dataset's authors or to Elsevier Research Data Support. Instead, please sign in using the Scopus Author Feedback Wizard to submit such requests, so the corrected data can be used in future annual updates of the citation indicator database.
Requests to correct your own Scopus data should go to Scopus directly through the Author Feedback Wizard (see above). Do not send these requests to Elsevier Customer Service or to the creators of the datasets (John Ioannidis or Jeroen Baas). Neither Elsevier Customer Service nor the creators of the datasets can amend the published datasets, since they are in archival form. Elsevier Customer Service can only help with questions about accessing or navigating the dataset and general platform support. The creators of the datasets should be contacted for methodological and technical issues that are not covered by the FAQ and the detailed methodological references (see below) and for potential research collaborations.
Please read the 3 associated PLoS Biology papers that explain the development, validation and use of the composite citation indicator metrics and databases
Ioannidis JPA, Boyack KW, Baas J. Updated science-wide author databases of standardized citation indicators. PLoS Biol. 2020 Oct 16;18(10):e3000918. https://doi.org/10.1371/journal.pbio.3000918
Ioannidis JPA, Baas J, Klavans R, Boyack KW. A standardized citation metrics author database annotated for scientific field. PLoS Biol. 2019 Aug 12;17(8):e3000384. https://doi.org/10.1371/journal.pbio.3000384
Ioannidis JP, Klavans R, Boyack KW. Multiple Citation Indicators and Their Composite across Scientific Disciplines. PLoS Biol. 2016 Jul 1;14(7):e1002501. https://doi.org/10.1371/journal.pbio.1002501
What do the listed retraction data mean?
Since 2024, the annual editions of the databases include also information for each top-cited scientist on the number of retractions, the total number of citations received to their retracted papers, and the total number of citations they have received from any retracted papers. The Methods and some analyses of the retraction data can be found in: Ioannidis JPA, Pezzullo AM, Cristiano A, Boccia S, Baas J. Linking citation and retraction data reveals the demographics of scientific retractions among highly cited authors. PLoS Biology 2025, 23(1):e3002999. https://doi.org/10.1371/journal.pbio.3002999 . See also further analysis in Boccia S, et al. Gender imbalances of retraction prevalence among highly cited authors and among all authors. bioRxiv 2025.07.08.662297; doi: https://doi.org/10.1101/2025.07.08.662297. Research Integrity and Peer Review (in press).
How can I correct errors in the databases?
The databases use Scopus data without further alteration. Only Scopus can correct their data and, if corrected, the correct versions will be used in the next iteration of our databases which are published annually in an archival form and will not change until the next annual update. The published version reflects Scopus author profiles at the time of calculation. We thus advise authors to ensure that their Scopus profiles are accurate. REQUESTS FOR CORRECTIONS OF THE SCOPUS DATA (INCLUDING CORRECTIONS IN AFFILIATIONS) SHOULD NOT BE SENT TO US. Instead sign in using the Scopus Author Feedback Wizard to submit such requests, so the corrected data can be used in future annual updates of the citation indicator databases. Similarly, if any data on retractions are deemed to be incorrect, requests for correction should be addressed to the Retraction Watch database.
Are the databases stratified by gender and age?
The main databases do not contain information on gender or age on the listed scientists. Information on publication age can be inferred from the listed year of starting to publish. However, one may peruse another database that we have created, and which provides information also on gender as well as stratified top-cited scientists according to publication age in 4 strata (first publication before 1992, 1992-2001, 2002-2011, 2012 or later). The database is in doi: 10.17632/wwykk8d48g See also: Ioannidis JPA, Boyack KW, Collins TA, Baas J. Gender imbalances among top-cited scientists across scientific disciplines over time through the analysis of nearly 5.8 million authors. PLoS Biol. 2023 Nov 21;21(11):e3002385. https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3002385
Does number of publications contribute to the composite citation indicator and respective ranking?
No, number of publications does not count at all. Impact may be achieved with very few or many publications. If a scientist with very few papers (even one or two papers) can achieve very high impact, they are not excluded from the database simply because they published few papers. Conversely, some scientists publish extremely high numbers of papers, and this may reflect either amazing, genuine productivity or manipulative practices. We have created a database that includes all scientists who have published more than sixty (60) full papers (full articles, reviews, or conference papers, excluding from the count all other types of published items) in any single calendar year. The database is in: https://elsevier.digitalcommonsdata.com/datasets/kmyvjk3xmd/2 See also: Ioannidis JPA, Klavans R, Boyack KW. Thousands of scientists publish a paper every five days. Nature. 2018 Sep;561(7722):167-169 https://www.nature.com/articles/d41586-018-06185-8 and Ioannidis, J.P.A., Collins, T.A. & Baas, J. Evolving patterns of extreme publishing behavior across science. Scientometrics (2024). https://doi.org/10.1007/s11192-024-05117-w
How could I correct affiliations?
To correct affiliations, please use the same exact process outlined under “How can I correct errors in the databases” sending the corrections directly to Scopus, not to us. We have no access or jurisdiction over the Scopus data. If corrected in Scopus, the corrected/preferred affiliation will appear in the next annual iteration of our databases. Scopus uses a machine learning approach to select only one affiliation from each author, based on the most recently published papers. This means that only one of possibly many affiliations is selected and this may not be the one that a scientist might prefer for themselves. It is impossible to know for millions of scientists which affiliation they prefer.
Does the database include deceased scientists, or is it limited to living individuals?
Both deceased and living scientists are included in the datasets. Scientists with very old publications start year are likely to be deceased, but also some young scientists may also be deceased. A recent/current year of latest publication does not offer complete proof that a scientist is alive, since for some deceased scientists some of their classic works may be republished currently and some scientists continue having their name listed in work they initiated when they were alive but is published posthumously.
How are self-citations calculated and what are considered to be too many self-citations?
Self-citations are counted from all authors of a given paper. E.g. if a paper X has 10 authors, all the papers that cite X and have one of the 10 original authors appear in their authors’ list are counted as self-citations. Percentiles of citations for each scientific discipline are provided in the databases so that one can place numbers in context of the discipline. For more details see: Ioannidis JPA, Boyack KW, Baas J. Updated science-wide author databases of standardized citation indicators. PLoS Biol. 2020 Oct 16;18(10):e3000918. https://doi.org/10.1371/journal.pbio.3000918
Can one infer evidence for gaming of citations from the database?
The presented data may offer some hints of citation orchestration and similar manipulations, but judgment needs to be cautious and linked with additional evidence. For potential metrics of citation orchestration, see: A. arXiv:2406.19219 [cs.DL] https://doi.org/10.48550/arXiv.2406.19219 B. Ioannidis JPA, Maniadis Z. In defense of quantitative metrics in researcher assessments. PLoS Biol. 2023;21(12):e3002408 C. Ioannidis JPA, Maniadis Z. Quantitative research assessment: using metrics against gamed metrics. Intern Emerg Med. 2024 Jan;19(1):39–47. D. Ioannidis JP. Features and impact in precocious citation impact. PLoS One. 2025 Aug 6;20(8):e0328531.
Do the data include only citations from original papers?
No, data are provided from all published items included in Scopus and these include large numbers of not only traditional articles, but also reviews, conference papers, editorials, notes, letters, news items, and more. Classification of article type may have substantial error in Scopus and the boundaries of what is “original” can be contested. E.g. some editorials may have some extremely original ideas, while some “original articles” may show zero originality. While most editorials get minimal or no citations, some may be highly cited, and this demonstrates impact for the writer(s). If a published work can be cited, citations should be counted. Nevertheless, for those interested in authors with extreme editorializing work, an analysis is provided in: Ioannidis JPA. Prolific non-research authors in high impact scientific journals: meta-research study. Scientometrics. 2023;128(5):3171-3184. doi: 10.1007/s11192-023-04687-5.
Why do some data look so different from Google Scholar and other citation databases?
Each citation database has its own inclusion and eligibility criteria and therefore different citations are counted. Our databases use Scopus data. In most disciplines, Google Scholar tends to count more citations than Scopus, as it is far less selective regarding the citing sources that it considers eligible. Many works have been published around issues with indexation, see for instance: Sauvayre R. Types of Errors Hiding in Google Scholar Data. J Med Internet Res. 2022 May 27;24(5):e28354. doi: 10.2196/28354. PMID: 35622395; PMCID: PMC9187964.
Why are others and not me included even though they have less?
We often get these questions by many hundreds of authors, and they practically always reflect a poor understanding of how the composite score is calculated (see above), that it is based on Scopus data, and that selection is subfield-adjusted. E.g. a selected scientist may have fewer publications than a non-selected one (publication counts does not count in the composite score calculation); fewer total citations than a non-selected one (many citation indicators are synthesized in the composite score trying to correct for co-authorship and relative contributions); or even lower all-science ranking than a non-selected one (the selected scientist may be ranked 1301000th across all-science, but may still be in the top-2% of their subfield, if this is a subfield that receives generally low numbers of citations, while a non-selected scientist may be ranked 104000th across all-science, but may not be within the top-2% of their subfield, if this is a subfield that receives generally very high numbers of citations).
Why is my listed primary subfield of research not what I would have described for myself?
The subfields are named using the Science-Metrix classification and nomenclature and the 174 subfield names may not precisely correspond to what specific scientists think about themselves. One should nevertheless examine in Science-Metrix what areas are covered under each subfield – a perfect name representing equally all of them would be impossible. Moreover, the primary subfield is the one with the largest proportion of papers for a given scientist and some scientists are split across many diverse subfields. The databases provide also the secondary (second more frequent) subfield for each scientist.
Are all scientists listed in the databases excellent scientists and those not listed are not?
NO! If an author is not on the list, it is simply because the composite indicator value was not high enough to appear on the list. It does not mean that the author does not do good work. Similarly, some of the listed scientists may not be excellent and very few may even have serious, well-documented violations of research integrity in their record (as can be transparently gleaned from the data that we provide on retractions). Citation metrics have limitations.
No, this follows from the answer to question 13 above. We recognize that since we have made our datasets publicly available, many other scientists and non-scientists, administrators, institutions, organizations, or other entities may have used them. E.g., they may have summarized the information that we have provided, recast the datasets, disseminated them, or interpreted them in various ways. We trust that open public availability is worthwhile for transparency, but we should not be asked to endorse any of the multifarious ways that these datasets may be used or misused. Again, we advise cautious use of citation metrics and careful perusal of the methods that have been used to generate the metrics in our datasets. If any entity sells (for money) certificates or other tokens based on the data, this is inappropriate, since we have made the datasets available for free public use excluding commercial use.
Can the data be used to appraise and rank institutions?
Many institutions pay increasingly a lot of attention to these data and may disseminate information on how many of their scientists are included in these databases as a signal of prestige. They often depend on re-uses of the databases by entities that place inordinate emphasis on excellence signaling and PR aspects and we worry that these entities may have communicated the data in various misleading ways. We recommend that institutions that use these databases should read carefully our Frequently Asked Questions file and understand the meaning, strengths and limitations of the data. Moreover, we recommend adoption of more balanced and nuanced approaches that do not focus only on number of scientists with high citation impact but also adjust for institutional size and consider also research integrity proxies. Institutions may use the databases not only for proxies of excellence but also for indicators of integrity risk. For example, see: Science-wide mapping of institutions based on affiliated authors’ impact and research integrity proxies. John P.A. Ioannidis, Jeroen Baas, Roy Boverhof, Cyril Voyant. bioRxiv 2025.12.16.694665; doi: https://doi.org/10.64898/2025.12.16.694665 (in press in Royal Society Open Science).
Why do the databases not filter obvious anomalies and outliers?
Some scientists have suggested that the publicly released data should filter and exclude upfront highly cited authors who are obvious anomalies and outliers. Candidate exclusions might be authors with large proportion of self-citations, very high number of retractions, journalists and editorial staff who massively publish journalistic and editorial pieces, scientists who belong to long past historical eras, and so forth (see for example: https://link.springer.com/article/10.1007/s11192-026-05682-2). However, the exact threshold for activating each such filter would be entirely arbitrary. More importantly, one of the main purposes of the databases is to allow rigorous scientific investigation in detecting authors with anomalous profiles and outliers, as demonstrated in several references cited in replies to previous questions. Upfront filtering with any type of filters, would diminish this crucial value. Scientists who wish to use these databases for answering specific research questions may decide to use filters or selection criteria that are fit for purpose given the research questions at hand.
What if I have worked in multiple countries?
We have added a new set of variables that captures whether an author has published a substantial portion of their career papers from institution/organization affiliations located in more than one country and, if so, which are these multiple countries. For more details on the methods, thresholds and definitions, see the preprint: Ioannidis JP, Baas J. Mapping scientific cosmopolitanism: scientists with publication records substantially split across multiple countries. bioRxiv 2026. Country assignments for authors on papers are in that process enriched and normalized using data on organisational structures, beyond the countries listed explicitly in the paper. This can cause the countries listed in the dataset produced to differ from what is visible on Scopus directly and this does not mean the data in Scopus is wrong: the enrichment process also introduces few precision errors. The estimated error rate for the multiple countries inferential process that we have used is 4%.
Did we answer your question?
Recently viewed answers
Functionality disabled due to your cookie preferences