An Estimated 146,932 Fake Citations Entered the Scholarly Record in a Single Year
An audit of 111 million references estimates 146,932 fake citations entered the literature in 2025, and peer review is not catching them.

A team I worked with had "the high schooler" completely figured out. One persona, a first name, a stock-photo face, five or six traits on a slide, and a product being built to serve her. I asked what research she was based on and the answer was none. Nobody had talked to a student. She had been assembled out of the team's own reflection, which is why she mostly resembled the team.
So I went and did the research, and the one high schooler turned out to be at least three genuinely distinct users. There were the students already on a well-resourced path to a four-year college. There were the students with no such path at all. And in between sat a third group who wanted more than the road they could see and had nobody to show them how to get there. The largest of those three groups was also the most underserved, and it was the group that looked least like the people in the room. A majority of Americans still never finish a four-year degree – 42.8% of those aged 25 to 39 hold one, on the Census Bureau's 2024 figures – so most real high schoolers were never on the manicured college track the team had imagined in the first place.
A wrong persona doesn't stay wrong in one place. Once your picture of the user is off, every decision downstream inherits the error, and you can run experiment after experiment and pivot after pivot and never climb out, because the actual fault is sitting all the way back at the start. The experiments are honest. The analysis is careful. The conclusion is still wrong, and nothing inside the loop is capable of telling you so.
I keep coming back to that shape, because a study released this spring describes the same failure at the scale of the entire scientific literature.
How many fake citations are actually in the literature
Six researchers audited 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN and PubMed Central, and arrived at a conservative estimate of 146,932 hallucinated citations entering the record in 2025 alone. The paper is Zhao et al., LLM hallucinations in the wild, and it's an arXiv preprint that has not been peer-reviewed. It would be strange to build an argument about citation hygiene on a source whose status I'd fudged. Nature, which covered the work, says the same thing plainly.
An estimate is an estimate, and the authors are explicit that theirs is a conservative one. Their method shows why: as Phys.org described it, over 95% of references matched successfully, unmatched entries had their typing errors corrected until a match appeared, Google Scholar was used as a final check for obscure publications, and only faulty references appearing in material published after 2022 were counted at all. Each of those choices pushes the number down rather than up.
The per-repository rates are low, and they vary more than you'd expect. SSRN, a social-sciences working-paper server, came in highest at 1.91%, nearly five times any other major repository, followed by arXiv at 0.39%, PubMed Central at 0.27% and bioRxiv at 0.21%. Under 2% is not a broken literature and I don't want to say it is. What's worth your attention is the direction: a sharp rise from mid-2024 onward, especially pronounced in fields with rapid AI uptake and in manuscripts that carry the linguistic signatures of AI-assisted writing.
There's a figure I wanted for this post that isn't in it. A per-paper prevalence trend has been circulating in summaries of the study, climbing steeply across 2023, 2025 and early 2026, and it's a better hook than anything I've written above. I opened the PDF to confirm it and couldn't find it, so it's out. Which is a little awkward, because it was the best number in my draft.
Peer review is not catching them
Fabricated citations are surviving expert review, and the study measured that directly. The authors traced bioRxiv preprints containing unmatched references through to their published versions in PubMed Central and found that 85.3% of the hallucinations present in the preprint persisted into the published record. I checked that one in the paper itself rather than in a summary, since it's exactly the sort of number that gets repeated loosely.
Conference review looks the same. The AI-detection company GPTZero ran two investigations, finding 100 hallucinated citations in NeurIPS 2025 accepted papers and more than 50 in ICLR 2026. Both of those reached me through a practitioner write-up that credited the NeurIPS work to a different author entirely, and I only caught that by opening Zhao et al.'s reference list, where both turn out to be GPTZero's own reports. I haven't read those two reports at source, so take the counts as well-attested rather than as something I verified myself. Reviewers read the NeurIPS papers closely enough to vote for accepting them, and not one of them opened a reference that didn't exist.
Reviewers aren't checking reference lists because reference lists were never the part anyone thought needed checking. A citation was a pointer to something that obviously existed, and the reviewer's job was to argue with the argument. So the habit of passing a source along because someone credible cited it was always a small shortcut, and it has quietly turned into an unsafe one.
"I checked the source" has stopped being a complete answer
There's a simpler version of this problem that gets most of the airtime, and it's the one I've written about separately: AI research tools invent sources, which is where the fake citations come from. That version has a clean fix. You verify the citation. It resolves or it doesn't, and you're done.
This is the harder one. If hallucinated references are entering peer-reviewed journals and staying there, then the corpus you verify against now contains the same class of error you're verifying for. "I checked it against a published paper" stops being a complete answer, because the published paper may itself be carrying a reference to something nobody ever wrote.
Which is the persona problem again, at the scale of a whole field. The fault sits at the start of the chain, before any of the careful work begins, so no amount of care further down will surface it. You can be genuinely disciplined about your reading and still end up standing on a claim that traces back to nothing. That's epistemic debt in its purest form: the interest accrues silently, and the bill arrives on the day somebody finally walks the chain.
What checking a citation properly looks like now
Provenance depth is three questions instead of one. Does the reference resolve to a real paper? Does that paper actually say what the citing paper claims it says? And what is its status, meaning preprint or peer-reviewed, retracted or standing, primary research or somebody's summary of primary research? Most of us who cite research in client work stop after the first question, and until fairly recently that was a reasonable place to stop.
If you advise clients for a living, you already run this discipline somewhere else. When a client hands you a number for a board deck, you don't put it in the deck. You ask where it came from, who calculated it, and against what definition, because your name goes on the recommendation and theirs doesn't. The literature has now earned the identical treatment. It's the same skill you already have, pointed at a different object.
I studied religion and poetry as an undergrad, which is an odd résumé line for someone who spends part of every week chasing down whether a paper exists. It's also the most useful training I got. A large share of what you do in a religion department is source criticism: who wrote this, when, from what, and who copied it down wrong somewhere along the way.
The pressure runs the other direction, though, because nobody reads the footnotes. A deliverable's citations are the part every reader trusts and almost no reader opens, which is precisely why a fabricated one can sit there for years. And what it looks like in a client report is not subtle at all once somebody does check.
Building a corpus you own rather than inherit
The practical answer is unglamorous, and it's the thing I'd want in place whether or not anyone ever paid me to build it. Keep your own corpus. For every claim you might use, record the source, the date you accessed it, the verbatim sentence you're relying on, and the status of the thing it came from. Then reuse that across engagements instead of re-inheriting someone else's summary every time the topic comes back around. It converts verification from a recurring gamble into a fixed cost you pay once per source.
That's most of what Fieldway's Research and Intelligence work is underneath the deliverables: answer one hard question in a way the client can check, then keep the answer current as the world moves. The reason it's built that way is the Cheap Yes. A fast, confident, nicely formatted answer that nobody can trace is the cheapest thing in the world to produce right now, and it's worth nothing at all when it turns out to be wrong, which you discover later and at full price.
It's also what replaced knowledge as the moat. Knowing things stopped being scarce some time ago. Being able to show where you know it from is scarce now, and it gets scarcer every month the record fills up with pointers to papers that were never written.
Checking a citation is about as glamorous as it sounds. You open the paper, you find the page, you read the sentence, and almost every time nothing happens – which is the whole point, and also why hardly anybody does it. Every paper I lean on in something a client pays for is one I've actually opened, and I can tell you which page.
I'd rather hand a client a short list of sources I've read than a long one I haven't.
Email matthew@fieldway.org to discuss.
Sources
- Zhao et al., LLM hallucinations in the wild: Large-scale evidence from non-existent citations, arXiv:2605.07723 (preprint, not peer-reviewed)
- Nature, Hallucinated citations highest in social sciences preprints site
- Phys.org, AI-generated fake citations are flooding scientific literature across four major repositories
- GPTZero, GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers (cited as ref. 7 in Zhao et al.)
- GPTZero, GPTZero uncovers 50+ hallucinations in ICLR 2026 (cited as ref. 6 in Zhao et al.)
- US Census Bureau, Census Bureau Releases New Educational Attainment Data (2024 data)
- Breaking Deep Research: where LLM search agents fail (misattributes the NeurIPS investigation)
More from Consulting Operations

806 Repositories, One Causal Design: the Speed Was Temporary, the Complexity Wasn't
A Carnegie Mellon study of 806 repositories found the AI speed gain faded within two months while the complexity it left behind kept compounding.

31% More Pull Requests Are Now Merging With No Review at All
Faros AI found 31.3% more pull requests merging with no review at all, and the incident data says that skipped gate is costing more than anyone tracked.

The EU Just Made "A Human Looked At It" a Legal Category
EU AI Act Article 50 took effect on 2 August. Its human-review carve-out quietly turns accountability into a legal category, not just good practice.
Want help running a sharper practice?
The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.
See how this works for advisors