AI Slows Developers Down at First. The Ones Who Push Through Come Out Faster.

The same developers who were 19% slower with AI in 2025 measured faster a year later. Judging AI at month one measures the dip, not the destination.

6 min readBy Matthew Stublefield
A rectangular object with a light on it

In early 2025, a group of experienced open-source developers were measured working with and without AI tools. They predicted the tools would save them about 24% of their time. The tools made them 19% slower.

That study, from METR, was one of the most-cited pieces of AI skepticism last year, and for good reason: it's a controlled measurement, not a vibe, and the gap between what developers expected and what actually happened was enormous. People felt faster and were slower. That's a genuinely important finding.

Here's what got less attention. METR went back to the same developers a year later. This time the measurement flipped: an estimated speedup of roughly 18%. Same people, and nothing about the underlying tools changed in a way that explains a swing that large. What changed was that they'd spent a year learning how to actually use them.

Before anyone frames this as vindication, the honest caveat has to come first, because it's the interesting part.

The number is soft, and pretending otherwise would be dishonest

METR is refreshingly candid that the follow-up is a weak signal. The 18% speedup carries a confidence interval running from 38% faster all the way to 9% slower, and the researchers say plainly that selection effects make the size of the improvement unreliable. Newly recruited developers in the same work showed only about a 4% speedup, with a confidence interval that also crosses zero.

So no, this is not proof that AI now makes developers 18% faster. Anyone selling you that number as a clean fact is doing to the good news exactly what last year's takes did to the bad news – amputating the uncertainty to make the headline land. The direction is real and shows up across the data. The precise magnitude is mush. Both of those things are true, and a claim that respects the reader has to hold them together.

I'd rather ship the honest version. The tools appear to help more now than they did a year ago, for people who've climbed the learning curve, and we can't yet say by exactly how much.

Every real tool has a J-curve. You already know this.

Strip out the word "AI" and this is the least surprising finding in the world. Every capable tool makes you slower before it makes you faster. The first month on a new IDE, a new framework, a new keyboard layout, a new methodology – you're worse. You're translating, second-guessing, fighting muscle memory. Then, if the tool is any good and you stick with it, you cross over, and eventually you can't imagine working the old way.

Managers accept this everywhere except AI. Nobody rips out a new deployment pipeline three weeks in because velocity dipped during the switch. We understand, in that context, that we're looking at the cost of the transition, not the value of the destination. But hand a team AI coding tools, measure them at week four, see a slowdown, and the conclusion gets written as "AI doesn't work" instead of "we are currently standing in the dip that every tool has."

The METR data, read across both years, is a picture of the dip and the climb out of it. The early study caught experienced developers at their worst moment with an unfamiliar tool. The later one caught the same people further along. If you'd judged them in month one and cancelled the program, you'd have thrown away the part that pays.

This is a measurement problem before it's a tooling problem

Here's where it gets practical, and where leaders get it wrong. The question isn't "does AI make developers faster" – asked flatly, that question has no answer, because it depends entirely on where the person is on the curve, plus a pile of other conditions. Stanford's research across roughly 100,000 developers found AI's effectiveness swings hard with task complexity, codebase maturity, language popularity, and codebase size. DX's longitudinal study of 400-plus engineering organizations found that as AI usage rose about 65%, median pull-request throughput rose just under 8% – a real gain, and a fraction of the 3x or 10x that leaders were promised and are now holding teams to.

Put those together and the honest summary is: modest on average, wildly conditional, and improving with maturity. That's not a satisfying number for a board slide. It's the truth, and it changes what a leader should actually do.

You stop asking for a verdict and start instrumenting the curve. Measure your own team, on your own codebase, over time – not once, at the bottom of the J, where every new tool looks like a mistake. Track the trend, not the snapshot. Give the adoption enough runway to cross over before you rule on it, the way you'd give any other capability investment time to compound. And when someone hands you a single dramatic number, in either direction, treat it as a starting question rather than a conclusion, because a point estimate with a confidence interval that crosses zero is a conversation, not a fact.

Instrumenting the curve is less exotic than it sounds. Pick a few outcomes you actually care about – how long a change takes to reach production, defect rates, how much rework a change tends to generate – and watch them move over quarters, on your own team, not against a competitor's press release. Compare a cohort three months into adoption against where they started, and expect the first readings to look bad, because that's the dip every tool has. Hold your judgment until you have enough slope to mean something. The teams that do this well treat AI adoption like any other capability investment: they give it runway, they measure the trend, and they decide on the trend. And when the line does turn up, they resist the urge to annualize it into a heroic number for a board slide, because the honest artifact is the slope itself, with its uncertainty still attached.

There's a small obligation in here too. Developers are the ones living inside this dip, being asked to adopt tools that make them briefly slower while executives read them week-four numbers. Judging people at the bottom of their learning curve – and worse, judging the tools by how those people perform at their most disoriented – isn't rigor. It's impatience dressed as data.

The uncomfortable, useful version

The clean stories are both wrong. "AI makes developers slower" was a real study read without its sequel. "AI makes developers 18% faster" is a real study read without its confidence interval. The version that survives contact with all the data is quieter: AI coding productivity is a learning curve with a genuine dip and a real climb, the size of the climb is still uncertain, and where your team lands depends on your code, your tasks, and how long you let the adoption breathe.

That's less tweetable and far more useful, because it tells you to do the one thing the hot takes never do – answer the question for your own team, with your own measurement, over enough time to mean something. Measure the curve. Don't sentence the tool at the bottom of it.

Want help running a sharper practice?

The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.

See how this works for advisors