Research Agents Get Less Accurate the More They Read: 42% Worse From 2 Tool Calls to 150
Deep research agents keep link validity above 92% at every depth while fact-check accuracy falls 42%, so here's the order to verify a report in.

I was trying to work out whether coaching could pay my mortgage. Not "is there a market for coaching" in the abstract – whether it covers the mortgage and the groceries and the gas. So I went looking for return-on-investment figures, and Perplexity came back fast and confident: put a dollar into coaching and get back something north of 700%.
That's a great marketing line. Hire me as your coach and you'll see seven times your money back. It also didn't pass the sniff test, so I dug in.
The citation was real, and so was the number. The study behind it measured team-level and organization-level productivity gains from developing leaders inside large enterprises, which is a legitimate finding and a useful one – improving the leadership bench at a big company genuinely can return in that range. None of it transferred to a B2C coach. Nobody paying me a thousand dollars was going to make an extra seven times as much the following week, and on inspection the framing implied something stranger still: that somebody earning a six-figure income who hired a coach would watch their salary climb 700%. What the tool had done was lift one sentence out of a study, extrapolate from it, and invent the surrounding context. It never dug far enough to notice the context didn't carry.
The error isn't the part I keep coming back to. What I keep coming back to is that the error was flattery. It told me exactly what I wanted to hear at the exact moment I most wanted to hear it, and that is the condition under which a person is least likely to check. Had it come back with "coaching ROI is hard to establish and most of the literature is enterprise leadership development," I'd have opened every source, because a disappointing answer makes you want a second opinion. A thrilling answer makes you want to go build the landing page.
I spent about a year on that question. Every useful thing I learned in that year came from opening a source and reading who the study was actually about, and every dead end came from trusting somebody's summary of it. So I dig, and dig, and dig, which is less a talent than a method I've never found a way to skip.
There's now a small pile of evaluation work on exactly what I walked into, and the shape of the finding is stranger and more useful than "AI makes things up."
There are two kinds of citation failure, and only one of them dies when you click it
A fabricated citation is the cheap failure. The author doesn't exist, or the title doesn't, or the URL 404s. You click it, it dies, you delete the sentence. It's embarrassing if it ships, but catching it costs one click and no domain knowledge.
The other kind is a statement hallucination, and it's a different animal. The source is real. The link works. The page is genuinely about the topic at hand. What's wrong is the relationship between the sentence in the report and the content of the page – the source doesn't support the claim, or supports a narrower version of it, or supports it for a population that isn't yours. That one survives clicking. It survives a link-checker. It survives a reviewer who spot-checks three citations and finds all three resolve. And because it looks verified, it gets quoted, pasted into a deck, and repeated by the next person, who now has your document as their source. That's debt you can't see in the output.
My 700% figure was the second kind. Every surface signal was clean.
The published accuracy numbers, and the task they came from
The most careful public look at this is a peer-reviewed review in JMIR, which reports Keplinger et al.'s test of five deep research agents on a dermatology literature review. That task is close to worst-case for these tools – a specialized, citation-dense domain where hallucination rates are known to climb – so read what follows as a stress test rather than as a general property of the products. On that task, Gemini 2.5 Pro Deep Research and Perplexity AI Deep Research produced roughly 47% to 50% of their references with fake authors or titles. OpenAI Deep Research did much better on the mechanics – about 95% of its citations were identifiable and about 70% completely correct.
The mechanics were the good news. Over 50% of OpenAI Deep Research's cited statements contained at least one subtle inaccuracy or misrepresentation of the source: results misinterpreted, findings cited out of context, details attributed to papers that didn't contain them. The tool with the cleanest citation list had the majority of its sentences drifting from what its sources said. Cheap failures down, expensive failures up.
A separate preprint, Detecting and Correcting Reference Hallucinations in Commercial LLMs, comes at it from the URL side and finds 3% to 13% of citation URLs hallucinated and 5% to 18% non-resolving across commercial systems, with deep research agents showing the highest rates despite generating far more citations. Gemini's deep research mode was highest at 13.3% hallucinated while producing 113.1 URLs per query. And DeepTRACE, from Salesforce AI Research, summarised here, measured citation accuracy between 40% and 80% across systems that still produced large fractions of statements unsupported by their own listed sources – I reached that one through a secondary summary rather than the paper, so weight it accordingly.
Most of what's written about AI research tools ranks them by feature list and price and never once tests whether the output is true. That's the thing I'd fault the genre for, and it's why you can read ten round-ups and still not know which tool will hand you a real citation attached to a claim it doesn't make.
More depth degrades the synthesis, not the sourcing
The result that reorganized how I use these tools comes from a preprint evaluation framework, Parsing and Evaluating Source Attribution in LLM Deep Research, and specifically from its ablation on research depth. Fact-check accuracy drops by approximately 42% on average across two frontier models as tool calls scale from 2 to 150. More retrieval produced less accurate citation, not more.
In the same ablation, link validity and content relevance both stay above 92% at every search depth. The authors put it plainly: the degradation is specific to factual synthesis rather than source selection. Depth doesn't make the agent pick worse sources. It keeps picking good, working, on-topic sources and gets progressively worse at saying true things about them.
The variance between models is part of the finding. GPT-5.4 showed the steepest decline, from 79% down to 17%, while Claude Opus 4.6 was the most resilient at 80% down to 58%. That's not "GPT is bad" – it's two models in one ablation, and the spread tells you the behavior isn't a fixed property of the category. The broader pattern in the same study is the same asymmetry: frontier models hold link validity above 94% and relevance above 80%, and land at 39% to 77% factual accuracy.
The URL preprint offers a mechanism that fits: multi-step retrieval and synthesis may amplify fabrication by compounding errors across retrieval steps. Every hop inherits the previous hop's framing, and nothing in the loop goes back to ask whether step three still means what step one meant. Which is, more or less, the retrieval layer doing what it was built to do.
Both of those are preprints, not settled science. The JMIR piece is peer-reviewed. I'd treat the direction as well-supported and the exact magnitudes as provisional.
The order to verify in
Most review habits check the cheap failure first, because clicking links is easy and feels like diligence. That order is backwards, and the depth finding says why: the signals you can check in a second are the ones that stay clean while the thing that matters degrades.
- Sort claims by what they're deciding. A twenty-page report might have four sentences a decision actually rests on. Verify those completely and stop pretending you'll verify the rest.
- Read the source for entailment, not existence. The question isn't "does this page exist" but "does the sentence in my report follow from what's on this page." Check the population, the timeframe, and the unit. My 700% number failed on population alone.
- Flag the claims that flatter you. Anything that confirms the answer you were hoping for goes to the top of the verify queue. This is the step nobody builds, and it's the one I'd have missed entirely.
- Run several narrow queries instead of one exhaustive one. If accuracy decays with depth, then five focused runs at shallow depth beat one 150-call marathon, and you get five separately checkable outputs instead of one synthesis you can't unpick.
- Check link validity last, and automate it. It's the cheapest signal and the least informative. Let a script do it.
None of that is exotic. It's the reading practice any senior advisor already applies to a junior analyst's memo, pointed at a new author. You already know how to ask "where did this number come from and does it mean what you think it means" – the shift is that you now have to ask it of a document whose surface quality is far better than a junior's, and whose confidence is uniform regardless of whether it's right.
When a research agent is the right tool
Use these tools for the job they measurably do well, which is finding things. Link validity and topical relevance hold up at every depth in the evaluation, and that's real capability – an agent will surface in one sitting what used to take me the better part of a day, and it will find several sources I wouldn't have thought to look for. The last mile is where it fails, and the last mile is the part with your name on it.
So the division of labor I've settled on is that the agent builds the source pile and I own every sentence that claims a source says something. That's slower than the demo suggests and much faster than doing it all by hand. It also means the artifact I keep isn't the report – it's the sources, read, with my own note about what each one actually establishes. When a claim gets challenged six months later, that note is the only thing that saves me. The firms that skipped it are already on the record – what this looks like in a client deliverable, and then what happens when those citations get published and other people start building on top of them.
I don't think the answer here is to trust the tools less. It's to move where the trust sits: from the report to the reading. The agent can tell you where to look, and it turns out that's a genuinely valuable thing to be told. Deciding what's there is still the job, and that job got more important the moment the output started looking finished.
Always search for reality. You have to know what's actually happening, and the only way I've found to know that is to dig in.
If you're wrestling with how to make AI-assisted research trustworthy enough to sign your name to, email matthew@fieldway.org to discuss.
Sources
- Parsing and Evaluating Source Attribution in LLM Deep Research, arXiv:2605.06635 (preprint)
- JMIR (2026), Major Breakthrough or Incremental Progress for Medical AI?, reporting Keplinger et al. (peer-reviewed)
- Detecting and Correcting Reference Hallucinations in Commercial LLMs, arXiv:2604.03173 (preprint)
- Breaking Deep Research: Where Retail User LLM Search Agents Fail (secondary summary of DeepTRACE, Tow Center, BBC/EBU)
More from Consulting Operations

'MBB Depth Without the MBB Team' Is Half an Answer
A wave of vendors offers boutique advisors MBB-depth research without the MBB team. It's real, it's cheap, and it's only half of what the work needs.

How to Research an Audience You've Never Served
Researching an audience you've never served is a different job: the whole task is telling real demand from polite interest. Here's how it's done.

Your Features Shipped. Your Metrics Didn't Move. The Gap Has a Name.
Your team ships on schedule and the numbers stay flat. That's the build trap — and cheap AI building made it easier to fall into and costlier to ignore.
Want help running a sharper practice?
The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.
See how this works for advisors