AI Helps Weak Teams 4x More Than Strong Ones. That's Not a Win.
AI cut lead time nearly 50% for low-performing teams and 10–15% for elite ones. That gap isn't the good news everyone is reading it as.

I've roughly tripled a team's productivity about ten times in the last decade, across different companies and industries. Not one of those was a speed problem.
A statistic has been circulating since the spring that seems to argue the opposite, and it's being read almost exactly backwards.
Plandek's 2026 Engineering Productivity Benchmarks, built on delivery data from more than 2,000 engineering teams, found that lower-performing teams using AI cut their Lead Time to Value by nearly 50% compared to similar teams without it. Top-performing teams improved by 10 to 15%. That's roughly a fourfold difference, and it landed as a feel-good headline: AI is the great equalizer, the struggling teams gain the most, the gap is closing.
The gap is not closing. And the clearest evidence for that is sitting in the same report.
What Plandek actually concluded
Read past the headline finding to the report's own summary and you find this: AI does not fix broken flow, and systemic constraints still determine long-term performance. Alongside it: low-performing teams remain trapped in rework, technical debt, and unplanned work.
That's not a caveat tucked in a footnote. It's listed as a headline finding, right next to the 4x number, and it says almost the opposite of how the 4x number gets quoted.
The rest of the data agrees. After AI, bottom-quartile teams still take more than 35 hours on average to merge a pull request while top performers merge in under 21. Top teams still ship to production in under 22.5 days; the bottom quartile still takes more than 62. Those are the same teams that just got a 50% lead-time improvement. The distance between them didn't close – it persisted, at roughly 3x, on the measure that actually reaches a user.
One more thing about this study, because it's been quietly misdated in a lot of the coverage: it was published in March 2026, drawn from anonymized Q4 2025 data. It's not a fresh July finding. It's a spring report that got a second life.
Three organizations, one conclusion
Google Cloud's DORA research reached the same place from a different dataset. Their 2026 work found that higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability. More change, moving faster, breaking more often. The 2025 DORA report's headline framing, as summarized here, was blunter still: AI doesn't fix a team, it amplifies what's already there.
Then ClearRoute, working from four years of enterprise delivery assessments, put it in almost identical words: AI amplifies existing delivery systems. Mature platforms accelerate. Manual processes, fragmented tooling, and governance bottlenecks create more work, faster.
So the three datasets don't disagree. Plandek measured lead time and found a big gain for weak teams. DORA measured throughput and stability together and found the second one moving the wrong way. ClearRoute measured the whole path to production and found the constraints untouched. Different dependent variables, one mechanism.
Amplification is the finding. The 4x is what amplification looks like when you point it at the single stage AI actually touches.
Why a faster weak team is a worse problem
A low-performing team is not low-performing because its engineers type slowly. In my engagements, it's rework, unclear direction, decisions nobody will make, work that shouldn't have been started, and handoffs where context evaporates. Those are the things that made them slow. AI addresses none of them.
What AI does is let that team produce the wrong thing, or the unclear thing, or the thing nobody decided on, considerably faster. The rework rate doesn't improve. The volume feeding it does. A struggling team that just got 50% faster at authoring is now generating unresolved problems at twice the rate, pushing them into a review and release process that was already the constraint, and the lead-time number on the dashboard goes green while the underlying condition gets measurably harder to dig out of.
And this is where Goodhart's Law does its usual work. When a measure becomes a target, it stops being a good measure. Lead Time to Value was a decent proxy for health back when it was expensive to move – you couldn't fake it, because getting faster required actually fixing something. AI made it cheap to move that number without fixing anything. The measure survived; what it measured did not.
Metrics are there to test whether your hypothesis about the system was right. They aren't there to prove your investment worked, and they were never there to grade your people. If your delivery metric improved and you can't name which constraint you removed, you haven't learned anything about your system – you've learned that the metric was easier to move than you thought.
What actually fixes a weak team
Early in my career I got promoted into a role and my boss moved all the worst employees onto my team. They'd been under a manager who couldn't hold anyone accountable, and the reputation had followed them.
Fixing that took years, and none of it was about speed. Of the group I inherited, I had to fire one, one retired, one transferred out, and one – the person everybody hated working with, a constant complainer who dragged his feet – turned into a star who brought me some of the best ideas I got that year. With the student workers I raised the bar hard, because I could see the budget cuts coming if we stayed a group whose job was essentially to be present and make sure nothing got stolen. The reset was brutal: somewhere around 92% turnover in a single semester. Afterward I got to hire people who actually wanted to be there.
Three years in, the team was winning awards, delivering enough new services that our budget went up, and local employers had started treating us as a hiring pipeline. Retention went from losing about two people every six weeks to people leaving only when they graduated – freshmen who joined and stayed four, five, six years.
There is no version of that story where the fix was throughput. It was accountability, a clear bar, the willingness to have conversations I didn't want to have, and enough time for a system to change shape. A tool that made those people 50% faster in year one would have produced more of the thing that wasn't working.
That's what the Plandek finding actually says. AI gave the weakest teams the largest gain on the one dimension that was never their actual problem.
The question to ask instead
If your lead time improved this year, the useful question isn't how much. It's which constraint moved.
Pull the numbers that AI can't flatter. Change failure rate. Rework as a share of capacity. Unplanned work. Time from merge to production, separately from time to merge. Defect escape rate. If throughput improved and every one of those held steady or got worse, you have amplification, and the direction it's amplifying in is not the one on the slide.
Then go find out why the team was slow in the first place, which is nearly always a question about decisions rather than typing. A team that can't tell you why it's building something doesn't have a velocity problem, and it will not acquire one by getting faster.
Every dataset published this year agrees that AI makes teams more of what they already are. That's genuinely good news if you've done the work. If you haven't, you just bought a bigger engine for a car with a bent axle.
More from Consulting Operations

Not One of 300 QA Engineers Fully Trusts AI-Generated Code
Three hundred QA engineers rated their trust in AI-generated code. Not one gave it full marks, and their workload rose while their headcount didn't.

Most M&A Happens in Private Markets. Most AI Was Trained on Public Data.
AI due-diligence tools mostly learned from public disclosures. Most M&A deals never generate a public disclosure at all.

82% of Clients Expect AI-Powered Service. Only 1 in 4 Firms That Built One See a Return.
A survey of 200+ accounting firms found 82% feel more AI pressure from clients. Only 1 in 4 who acted on it saw a real return.
Want help running a sharper practice?
The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.
See how this works for advisors