How we investigated ghost references in academic articles
Joris Veerbeek
Earlier this year, dozens of accepted papers at NeurIPS, one of the world’s most prestigious AI conferences, were found to cite scientific publications that did not exist. This despite the fact that every manuscript passes through a rigorous peer-review process and only a fraction of submitted papers are accepted. In the months that followed, similar cases appeared at other scientific conferences and in academic journals. For De Groene Amsterdammer and Data School (Utrecht University), this prompted a closer look at Dutch academia. How often do publications by researchers at Dutch universities contain such ghost references? And is their number increasing as generative AI becomes more deeply embedded in academic practice?
To investigate this, we first compiled as complete an overview as possible of recent Dutch academic output. For this, we used OpenAlex, an open database containing hundreds of millions of research publications and offering broader coverage than commercial alternatives. Through the OpenAlex API—a technical interface that provides access to the database—we retrieved all open-access publications published between 1 January 2023, shortly after the introduction of ChatGPT, and 23 February 2026, for which at least one author was affiliated with one of the fourteen Dutch universities represented by Universities of the Netherlands, the umbrella organisation for Dutch universities.
This initially yielded 196,717 unique publications. We then attempted to extract the full-text PDFs for each of these publications. We succeeded for 146,264 articles, approximately 75 percent of the total. The remaining publications were excluded from the analysis because no complete text could be obtained. These included missing or faulty PDF files, records that had been incorrectly classified as scholarly articles in OpenAlex, articles lacking unique identifiers, publications that remained behind paywalls, and technical access restrictions that prevented automated downloads.
Next, we used GROBID to extract the reference lists from the PDFs. This software automatically identifies and structures references in scholarly articles. We excluded articles without references, as well as those from which no references could be extracted, from our further analysis. In the end, we analysed 136,471 articles. We then verified in four steps whether the scientific references in these publications actually pointed to existing scholarly works.
First, whenever available, we verified the reference’s DOI, a unique and persistent digital identifier assigned to scholarly publications. We did so by programmatically querying the DOI through Crossref. At this stage, we checked only whether the DOI itself existed; any discrepancies between the DOI and the remainder of the reference were therefore not taken into account.
Second, if the DOI returned no result, or if no DOI was provided, we searched for the reference title in four academic databases in succession: Crossref, OpenAlex, Scopus, and Google Scholar. From each database, we retrieved the five highest-ranked results. To allow for minor differences, such as spelling mistakes or variations in punctuation, we compared the retrieved titles with the original title using a so-called fuzzy matching method implemented through the Python library RapidFuzz. When the similarity score was at least 90 percent, we considered the result a match.
Third, when a reference could not be located in any of the academic databases, we used the open-source language model GPT-OSS-20B, with access to the DuckDuckGo search engine, as an additional search tool. This allowed us to identify references that had been missed in the previous steps or to detect possible errors in the PDF extraction process, such as footnotes that had been mistakenly classified as references. The language model could therefore only rule out a reference as problematic; it was never used to label a reference as a ghost citation.
If the third step likewise produced no result, we carried out a final round of extensive manual verification. To keep the number of manual checks manageable and focus on the most “severe” cases, we restricted this process to articles for which at least two references could not be identified in the previous steps. We first searched for the publication using standard search engines. We then checked whether the reference actually appeared at the indicated location by verifying its bibliographic details, such as journal, volume, issue, and page numbers. If the reference could not be found there, or if those details corresponded to an entirely different publication, we classified it as a ghost citation. In our study, only references that were manually verified in this way were counted as ghost citations. Since we limited manual verification to articles containing at least two untraceable references, the actual number of articles with ghost citations is likely to be higher.
Our definition of a ghost citation relied on two criteria. First, the title itself had to be untraceable: we could not find it anywhere online or in academic databases. Second, the bibliographic information—such as the journal, volume, issue, and page numbers—could not lead us to the publication in question. Only when both checks failed to identify an existing source did we classify a reference as a ghost citation.
The cases therefore did not involve simple typographical errors or minor inconsistencies in citation formatting, but combinations of details that could not possibly refer to a single existing publication. In some instances, the authors, journal, volume, and page numbers were all correct, while only the title differed substantially. We initially recorded such cases as uncertain, but ultimately excluded them from our analysis. When we contacted several of the researchers involved, other—often technical—issues turned out to be responsible, including errors in bibliographic management software. We took a different view of references that included the names of real authors but whose titles, journals, and other bibliographic details did not correspond to any existing publication. These were counted as ghost citations, since the presence of a genuine author name alone does not demonstrate that the cited work itself exists.
While our study was already underway, both Nature and The Lancet published investigations into the presence of ghost citations in the scientific literature. Our approach can, in some respects, be seen as a combination of the two. Like the researchers behind The Lancet study, we systematically checked references across multiple academic databases and likewise used a language model to screen potentially problematic citations. Our extensive manual verification process, in turn, closely resembles the approach described by Nature.
To combat ghost citations, researchers are now actively searching for a method that can detect them fully automatically. Based on our findings, however, we remain somewhat cautious. The overwhelming majority of references flagged by our system because they could not be retrieved automatically turned out to have legitimate explanations. These included publications from the pre-internet era, grey literature, translated references, contributions to conference proceedings that had not been indexed individually by search engines, articles listed as “submitted” or “forthcoming,” minor errors in titles, and references to book chapters behind paywalls.
Of the references we examined manually, fewer than one in ten ultimately turned out to be genuine ghost citations. This raises the question of whether a fully automated system can reliably detect the problem. Although it is entirely possible that other systems perform more accurately than ours, the distinction between ghost citations and legitimate but difficult-to-trace sources is often highly context-dependent, making it likely that false positives can never be eliminated completely.
In total, we identified 208 articles by researchers affiliated with Dutch universities containing a combined 748 references to non-existent studies. These articles involved 443 individual researchers connected to Dutch universities. For these figures, we restricted our analysis exclusively to publications that appeared in peer-reviewed journals. We therefore manually verified whether preprints and publications of unclear provenance had ultimately been published in a peer-reviewed outlet. Although we also encountered a small number of doctoral dissertations in our dataset, we did not include them here. We likewise excluded two articles featuring the fictional scholar “Dean Sackker” and a host of fabricated references, which were explicitly intended as a critique of the academic system.
We also linked journals, through their ISSNs—a unique identifier used to distinguish serial publications worldwide—to the Norwegian Kanalregisteret, where independent disciplinary committees classify scholarly publication channels according to quality level: Level 2 (internationally leading journals), Level 1 (recognised scholarly journals), Level 0 (not recognised as a scholarly publication channel), and Level X (journals whose scholarly status is disputed). This shows that more than two-thirds of the articles containing ghost citations (72%) were published in recognised scholarly journals (Levels 1 and 2).
Measured against the more than one hundred thousand scientific articles in our dataset, which covers the period from 2023 onwards, the number of articles containing ghost citations remains small: roughly one in every 550 articles includes fabricated references. The vast majority of the scientific literature can still be traced back to existing sources. Yet the number of ghost citations is rising rapidly. From 2025 onwards in particular, the share of articles containing ghost citations increases sharply; in the most recent period, it already amounts to one in every 200 articles. This trend is likely only becoming visible now because months, and sometimes years, often pass between the writing, review, and publication of scientific papers.
As these references increasingly find their way into academic databases, the question is how the tide can still be turned. More than a year ago, over 1,800 scholars signed an open letter titled “Stop the Uncritical Adoption of AI Technologies in Academia.” The letter focused primarily on the ways in which generative AI may undermine students’ ability to think and write independently, and even called for a general ban on its use in student assignments. More broadly, public debate about AI in universities continues to revolve largely around inappropriate use by students. Our findings suggest that this focus is too narrow. It is not only students, but also researchers further up the academic hierarchy, who are deploying generative AI in ways that raise important questions about scientific integrity, academic norms, and institutional responsibility.