Most M&A Happens in Private Markets. Most AI Was Trained on Public Data.

AI due-diligence tools mostly learned from public disclosures. Most M&A deals never generate a public disclosure at all.

5 min readBy Matthew Stublefield
A black and white photo of a sign that says privacy please

At an IMAA Institute webinar on applied AI in M&A, James McVeigh – CEO of Cyndx, an AI-driven deal-sourcing platform – put the core tension in one sentence: most AI systems are trained on public data, while most M&A activity happens in the private markets. I want to be upfront that McVeigh runs a company whose product is built specifically to address that gap, so his framing isn't disinterested research. It's also, in my read, correct – and worth taking seriously precisely because the incentive to say it doesn't make it wrong.

The mismatch, and what actually backs it

Here's what I could independently confirm, and what I couldn't. Third Bridge, which serves private equity research teams and has no stake in Cyndx's narrative, states plainly that generic AI models trained on public content "cannot replicate insight derived from proprietary expert calls, operator interviews, structured deal data, and internal diligence materials." DFIN confirms the input side of the problem: deal-sourcing AI commonly pulls "from a wide range of sources, including SEC filings, databases and social media" – rich terrain for a public company, thin-to-nonexistent for a private one. And Stambaugh Ness, writing about M&A research more broadly rather than AI specifically, makes a related point that's easy to miss: even academic studies of M&A success skew almost entirely toward public-company deals, because private company data is, definitionally, private – "difficult, if not impossible" to collect at the scale research requires. The bias McVeigh is describing in AI tools isn't new to AI. It's the same bias that's shaped M&A scholarship for decades, now baked into the training data of the tools built to analyze it.

What I could not independently confirm is a clean percentage. I looked for a published split of M&A deal volume between private and public targets from the data houses that would normally carry that number – PitchBook, Bain, McKinsey, S&P Global, Dealogic – and none of them publish it as a clean share. The closest proxy I found, from deal-sourcing vendor Grata, puts private-market activity at roughly 30,000 to 50,000 transactions a year globally, around 200 a day – a real scale figure, but an absolute count, not a percentage, and not paired anywhere with a comparable public-deal number that would let you turn it into a share. So "most M&A happens in private markets" is the industry's near-universal working assumption, echoed by everyone I read on this, and I couldn't find a single source that quantifies it as a clean statistic. I'd rather tell you that plainly than manufacture a percentage that sounds more precise than what actually exists.

Why the gap matters more than it sounds

If you're only using AI tools to sift public filings and news for a public-company deal, this mismatch barely touches you. The moment a private-market target enters the picture – which, on the working assumption above, is most of the time – the tool's training data thins out exactly where the deal risk concentrates. Structured deal data, operator interviews, the texture of how a founder actually runs the business day to day: none of that lives in SEC filings, because none of it was ever required to. An AI tool built mainly on public disclosure isn't wrong when it evaluates a private target. It's confidently answering with whatever public-adjacent proxy it can find, and a confident wrong answer is a worse diligence failure mode than an obviously incomplete one, because it doesn't prompt anyone to go look further.

Cyndx's own pitch at the same webinar is a useful signal here, even setting aside McVeigh's incentive to make it. The platform claims coverage of 33 million companies across more than 100 languages, with particular emphasis on private companies that wouldn't show up in English-language public databases. I can't independently verify that figure – it's vendor marketing from the same talk, not an audited number – but the fact that a commercial product is being built specifically to plug this gap tells you the market believes the gap is real and worth paying to close. Vendors build toward wherever they think the pain actually is.

This is a different failure than the ones I've written about elsewhere in this cluster. The diligence gap that kills nearly half of dead deals is about what buyers find late, once they're actually in the data room. Diligence timelines stretching without getting deeper is about where the added calendar time actually goes. Both are about the diligence process. This is about the diligence tooling itself carrying a structural blind spot before the process even starts – a distinct problem from either of those, and one I don't think most teams adopting AI-assisted diligence tools have separately accounted for.

It also sits right next to a related but separate risk I've covered: verifying whether a target company's own AI capability claims hold up. That's about catching a target that's overstating what its AI actually does. This is about something the diligence team brings into the room themselves – the tool doing the evaluating, not the company being evaluated, carrying the blind spot.

What I'd actually do differently

If a diligence process leans on general-purpose AI tools and the target is privately held – which, again, is the likely case most of the time – the practical move is treating the tool's output as a public-data-shaped first pass, not a complete picture, and being specific about which parts of the analysis it's actually equipped to do well. Public-facing risk factors, industry benchmarking against comparable public companies, sentiment and news-based signal – reasonable territory for a public-data-trained tool. Anything that depends on proprietary deal structure, direct operator input, or the kind of on-the-ground texture that never generates a filing – that's exactly where a human diligence process still has to do the work the tool structurally can't, no matter how capable the model underneath it gets.

Concretely, that means asking a version of one question before trusting any AI-assisted output on a private deal: what did this tool actually have to read to reach this conclusion, and was any of it proprietary to the target, or all of it adjacent public signal about comparable companies? If the honest answer is the latter, the output is a reasonable starting hypothesis and a bad final answer. Most of the risk in a private deal was never going to show up in the material an AI tool trained on public disclosure has access to in the first place – which means the gap McVeigh is describing isn't really a limitation of any particular model. It's a limitation of what "public" ever covered, and no amount of model improvement closes a gap that's about what data exists, not about how well anything reads it.

Want help running a sharper practice?

The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.

See how this works for advisors