KPMG Published a Report on Agentic AI. 40 of Its 45 Citations Were Fabricated.
I reported a bounce rate that was real and wrong, which is the small version of what just happened to four of the world's largest advisory firms.

A while back I told a client their bounce rate had come down a lot. I'd pulled the figure myself, out of their own analytics tool, and the query was right – the number was real, and it went into a deck with my name on it.
Then a few people who knew that site well looked at the slide and said, more or less, that's crazy low.
They were right, and the reason sat underneath the number rather than in it. The tool's bounce-rate definition blended two populations that behave nothing alike: cold inbound traffic from search, and warm traffic arriving from other pages on the same site. Search bounces high, because the visitor is cold and may have landed on the wrong page entirely. An internal referral bounces low, because that person is already there on purpose. The tool was doing exactly what it had been built to do, and the definition underneath the number was simply wrong for the claim I'd built on top of it. I hadn't looked.
I still have that deck in a folder somewhere. I've never gone back and fixed the slide, which probably says something about me.
What I took from it is narrow and it has held up. Verifying that a source exists and returns a number is not verification. You have to know how the number is constructed, because a real figure resting on a definition nobody inspected will read as verified right up until someone who knows the domain looks at it.
There is a much more expensive version of that failure working its way through the four largest professional services firms in the world right now.
What actually happened, firm by firm
All four of the largest professional services firms have now had AI errors surface in client-facing work, and they have responded four different ways. Three of the four involved citations that didn't survive being clicked. Most of the coverage blurs the distinctions between them.
The number in the headline belongs to KPMG. GPTZero's forensic review of KPMG's October 2025 report Total Experience: Redefining Excellence in the Age of Agentic AI found that of 45 cited sources, only 5 pointed correctly to real, uncorrupted sources; 28 contained paraphrased titles or fabricated components attached to real sources, 12 were too vague to verify at all, and 40 of the 45 citation titles were fabricated. KPMG pulled the report in June 2026, as reported by the Indian Express. Before that, UBS, the UK's National Health Service, Swiss Federal Railways and Transport for London had all said the report's claims about their AI usage were either untrue or misleading, according to the Financial Times. A KPMG International spokesperson said the report had been removed, that the firm was reviewing the circumstances of its publication, and that "we expect all our people to follow our guidelines on the responsible use of AI, including human oversight to validate content and verify independent sources."
The outcomes diverge from there. Earlier GPTZero investigations led EY and KPMG to withdraw reports. Deloitte's case was a different shape: it refunded part of its fee to the Australian government after AI-related errors were found in a report it produced under a government contract, per eWeek – a client engagement rather than published thought leadership, and errors rather than fabricated references specifically. PwC Middle East has not withdrawn anything. A PwC Middle East spokesperson said the firm "takes the accuracy of our published research seriously and is updating a limited number of supporting citations in the reports mentioned." Two retractions, one partial refund, one amendment in progress.
The PwC Middle East episode is separate, with its own numbers. GPTZero examined four PwC Middle East reports published between 2024 and 2026. One of them, Transforming Governance (2025), scored an 84% chance of being entirely AI-generated on GPTZero's own detector, rising to 100% when the reference section is excluded. That report promoted a framework called "Citizen Pulse," described as in use by the governments of Denmark, Saudi Arabia, the United States and Australia; GPTZero found no public evidence the framework existed outside the report at all, and concluded that PwC "appears to have hallucinated both an entire product and business dealings with four separate nations." A cited Riyadh air-quality study couldn't be located in the journal named, in any database, or under the listed authors, per eWeek. A JPMorgan AI success story cited in one of the reports traced back to a Medium post by a teenager with roughly 280 followers, describing a project that predates ChatGPT, reported by Yahoo Finance. And one cybersecurity report footnote still carried utm_source=chatgpt.com in the URL, documented by eWeek, which is the detail I keep coming back to, because it means somebody pasted a link out of a chat window and nobody opened it again on the way to a client.
GPTZero isn't a disinterested party here – it sells AI-detection tooling, and those detection percentages come out of its own product. What makes the findings usable anyway is that the Financial Times verified them independently, as reported here, and that the citation-level checks are the kind anybody can repeat with a browser, because either the study is in the journal or it isn't. GPTZero researcher Paul Esau told the Financial Times, in coverage carried here that the treatment of sources across the four PwC Middle East papers "was typical of AI-generated research."
None of the reporting supports an intent to deceive, and none of it says that all Big Four AI research is fabricated – the investigations cover specific named reports, not the firms' output as a whole. What it supports is that nobody clicked.
A fabricated citation passes review because it looks correct
GPTZero has a name for the pattern: it calls this "vibe citing", references that look authoritative until someone actually clicks them. The mechanism explains how these got past several rounds of review by people who were paying attention. A fabricated citation is structurally well-formed. Author surnames in the right position, a plausible journal, a year that fits the argument, a title that sounds like a title in that field. Everything a reviewer can check at a glance checks out. The one thing wrong with it is that the source doesn't exist, and existence isn't visible in the formatting.
I said something similar about AI-written code a while back: it produces plausible code, and plausible is harder to catch than bad. Reviewers were trained to spot mistakes, not to spot confidence. Same failure mode, different artifact. It's getting harder rather than easier as research tools read more on their own – I've written about why the tools do this – and at the volume they now operate at, the record itself is contaminated. "I found a source that says so" is a weaker claim than it was two years ago.
Most of what I read for pleasure is fantasy fiction, so I have a fairly high tolerance for invented sources. It's much lower in something a client is going to act on.
Why a large review pyramid didn't catch it
A committee is, most of the time, a machine for diluting accountability until no single person can be blamed for anything, and a review pyramid working to a publication date is a committee. The partner assumes the manager checked the sources, and the manager assumes whoever assembled the deck pulled them from somewhere real. That person assumes the research tool returned real things, because it returned things shaped like real things. Every layer added polish, and every layer assumed existence had been handled one layer down.
A review pyramid is built to grade whether a document looks finished rather than whether it's true, and that's a design choice rather than bad luck. Formatting, house style, tone, structure, legal exposure – there's a named owner for each of those. Existence has no owner, so it falls through.
KPMG's own statement names the missing piece, which is human oversight to validate content and verify independent sources. That phrase has been quietly turning into a legal term of art rather than a statement of good intent, and there's now a regulatory answer to what counts as human oversight.
Across my consulting work, maybe one leader in twenty feels a genuine obligation to the people on the other end of their decisions. That's my own count over twenty-some years, not a statistic, and the difference rarely looks like much from the outside. The one in twenty isn't more competent than the nineteen; they just treat the output as theirs. A reviewer with a checklist confirms that a citation is present. A reviewer who feels the obligation opens it, because their name is going on the thing and somebody is going to make a decision from it.
What a research buyer should ask for now
Verifiability is something you can put in a scope of work, and until this summer most buyers didn't think to ask for it.
Ask that every citation opens. Not "sources included" – opened, all of them, by a person, and that stated in writing. It's tedious and it costs a couple of hours, and it's now about the cheapest insurance in professional services.
Ask for a confidence level with the reasoning that produced it. "We're highly confident" is worth nothing standing alone. "We're highly confident because three independent sources agree, and here's how we reconciled the definitional difference between two of them" is something a stranger can check.
Ask whose name is on it. Not the firm's name – the person's. Somebody should be able to say out loud, I made this call, and if it's wrong that's on me and I'll fix it. That sentence is terrifying the first few times you say it, and it's the price of being trusted with a decision.
The money moved ahead of the buyers here. Consulting valuations took a beating on the theory that AI eats the research layer, and what the market already priced looks less like a panic every time one of these reports gets pulled. Firms overstating their own AI are learning separately when the claim becomes a liability rather than just an embarrassment.
What a verifiable process looks like in practice
Two rules came out of my bounce-rate mistake, and both are boring on purpose.
Cross-reference, always, and never report off one source. On analytics work that means the product analytics tool, the web analytics tool and the warehouse view, reconciled against each other before anything goes in a deck. Where two disagree, the disagreement is the finding, because it usually means a definition differs underneath – the exact thing that got me. Three sources is enough to notice; one is only enough to be confident.
Then publish a confidence level with the reasoning behind it, which is the part that lets someone else check you. The practice itself lives in a markdown file in my Obsidian vault, which is less of a system than that sentence makes it sound.
If you run a boutique advisory firm, or you've been independent for fifteen or twenty years, the third ask is already true of you and there's no getting out of it. Your name is on the deliverable. There's no layer beneath you to absorb a bad citation and no committee to spread it across. That has always been described as the constraint of working small, and this summer it started reading as the reason to hire you instead. The firms with the deepest review benches just made the argument on your behalf.
My number was real and wrong at the same time, and the only thing that has ever stood between those two is a person who went and looked underneath it.
Email matthew@fieldway.org if you want to talk through how to check research before your name goes on it.
Sources
- GPTZero, Chasing the Hallucinations: KPMG's AI-Powered Attempt at "Redefining Excellence"
- GPTZero, PwC report hallucinates product and government customers
- Financial Times, KPMG report contained AI hallucinations on benefits of . . . AI (paywalled)
- Indian Express, KPMG retracts agentic AI study after researchers flag hallucinations, fake citations
- eWeek, GPTZero Flags Suspected AI Content and False Citations in PwC Reports
- Yahoo Finance, PwC AI reports tainted by hallucination errors
- Yahoo Finance, PwC Gets Caught Passing AI Slop Off as Authentic Research
- Business Standard, PwC reports riddled with AI-generated errors, fake citations
More from Consulting Operations

'MBB Depth Without the MBB Team' Is Half an Answer
A wave of vendors offers boutique advisors MBB-depth research without the MBB team. It's real, it's cheap, and it's only half of what the work needs.

How to Research an Audience You've Never Served
Researching an audience you've never served is a different job: the whole task is telling real demand from polite interest. Here's how it's done.

Your Features Shipped. Your Metrics Didn't Move. The Gap Has a Name.
Your team ships on schedule and the numbers stay flat. That's the build trap — and cheap AI building made it easier to fall into and costlier to ignore.
Want help running a sharper practice?
The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.
See how this works for advisors