Your Delivery Dashboard Might Be Measuring Nothing Anymore

AI didn't break your delivery metrics. It made most of them meaningless – and the one you think is still safe might not be either.

5 min readBy Matthew Stublefield
Analytics dashboard for content marketing with 1981 Digital

Faros AI looked at 22,000 developers across 4,000 teams last year and found individual output climbing fast under AI adoption – 21% more tasks completed, 98% more pull requests merged per developer. Then it checked the metrics that are supposed to tell you whether delivery actually got better. They didn't move.

That gap is the whole story. If you run engineering, you almost certainly have a dashboard built on some version of the four DORA metrics: how often you deploy, how long a change takes from commit to production, how often a deployment fails, and how fast you recover when it does. Those four numbers have been the closest thing the industry has to an honest scoreboard for a decade. They were built for a world where more code took more human time to write. That world is gone, and nobody sent the dashboard a memo.

Two of your four metrics just started lying

Start with deployment frequency and cycle time, because the mechanism here is the least controversial part of this story. LeadDev's reporting on the metrics AI has scrambled makes the point plainly: AI has severed the link between effort and output, inflating deployment frequency, cycle time, and PR volume without improving quality. An agent can produce a "working" implementation in minutes. A trivial, AI-assisted change is nearly frictionless to ship. So deployment frequency goes up – but as LeadDev's sources put it, what it's measuring has quietly changed from engineering maturity to tool adoption.

Nicholas Arcolano, who leads AI and research at Jellyfish, told LeadDev that AI-generated code is running about 20% larger on average than human-written code for the same task. Nobody's published the dataset behind that yet, so treat it as a practitioner's read rather than settled research – but it rhymes with everything else in this piece. More lines, more commits, more merges. None of it is evidence anyone's system got healthier.

DORA's own research team – the people at Google Cloud who invented these four metrics in the first place – put it more bluntly than a vendor blog would dare: AI can easily inflate the volume of code generated, and leaders need to stop relying on narrow, output-based metrics as a proxy for real productivity. Their 2026 research found something that should unsettle anyone reading a dashboard that only goes up and to the right: higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability, at the same time. Faster shipping is arriving with more incidents, not fewer. That's not a footnote. That's the headline, and most teams are reading the chart upside down.

The metric everyone assumes is still safe

Here's where it gets interesting, and where I think most coverage of this topic has been too tidy. LeadDev's expert sources single out mean time to recovery – MTTR, how long it takes to fix things once they break – as the one holdout. Their argument is intuitive: MTTR depends on a real production incident and a real recovery. You can't fake that with a fast commit. It remains, in their words, one of the more reliable signals of engineering health.

I want that to be true, because it would be tidy. It isn't quite. DORA's own March 2026 research groups MTTR together with the other three metrics as all under significant pressure from AI adoption – a meaningfully more pessimistic read than "the one that still holds." Faros AI's analysis of the 2025 DORA data lists MTTR among the metrics that stayed flat alongside lead time, deployment frequency, and change failure rate, not as the exception. So you've got a practitioner-commentary outlet calling MTTR the trustworthy one, and the two sources with actual telemetry declining to back that up.

I don't think LeadDev's sources are wrong, exactly. I think they're describing what MTTR should still measure, and the research is describing what's actually happening to it. Those are different claims, and the gap between them is worth sitting with before you bet your quarterly review on it. If someone tells you MTTR is the one number on your dashboard AI hasn't touched, ask them which study they're citing. There's a good chance the answer is a vibe, not a dataset.

What to actually watch instead

None of this means throw out the dashboard. It means stop reading it the way you used to. A few adjustments, in order of how cheap they are to make:

Stop treating throughput and instability as separate conversations. DORA's finding – that both rise together under AI adoption – means a happy deployment-frequency chart next to a quiet incident channel isn't good news. It's a lagging indicator. Look at them on the same axis or don't look at either alone.

Discount any metric that's purely about volume: lines of code, PR count, commits per developer. Faros's numbers – 21% more tasks, 98% more PRs – describe individual activity climbing while organization-level delivery metrics sat still. Activity metrics were always a proxy for value delivered. AI broke the proxy without breaking the activity, which is exactly backwards from what you want a metric to do.

Watch what DORA calls the verification tax: their research found 30% of developers currently report little to no trust in the code AI generates for them, which means a real, hidden cost is going into review and validation that doesn't show up anywhere on a standard delivery dashboard. I've written before about what that trust gap does to your review queue – it's not a side issue, it's the tax collector for everything above.

And be honest about which claims in this space are research and which are marketing. Faros AI's telemetry is real and large, but it's a vendor's self-selected customer base, not a neutral survey – useful, not gospel. The same goes for any tool that shows up promising to fix your AI-era metrics problem. Speed that doesn't show up as delivered value is the pattern of this whole moment, and a new dashboard vendor is not automatically the exception to it.

The dashboard isn't lying. It's answering a question you stopped asking.

If your team's delivery metrics all look fine right now, that's not evidence you're fine. It might just mean nobody's asked the dashboard a hard question lately, and it's still faithfully answering the easy one. A number that used to mean "we're shipping well" now sometimes just means "we're shipping a lot," and those stopped being the same sentence the day AI started writing a fifth of your diffs.

The teams that get hurt by this won't be the ones with bad metrics. They'll be the ones with metrics that still look good, for reasons nobody checked.

Want help running a sharper practice?

The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.

See how this works for advisors