Your Metric Isn't a Target. It's a Hypothesis You Forgot to Test
Features ship and the metric stays flat because teams treat the metric as a target to hit, not the hypothesis it was built to test.

I once led a team whose success measure had quietly become, in practice, features shipped – the count of bullet points in the release notes. Nobody wrote that down as the goal. It just became the thing that got celebrated, so it became the thing people optimized. And the team was good, so they hit it. They stuffed the interface with capabilities nobody had asked for, the product got harder to use, and month after month the team hit its number while active users declined. Instrumentation later showed the majority of engineering effort had gone into features that were barely touched.
The people doing this weren't cynics gaming the system. That's the part worth sitting with. They corrupted the metric by taking it seriously – by treating a number that was supposed to test whether they were creating value as a target to hit instead. That's Goodhart's law in one sentence: when a measure becomes a target, it stops being a good measure. And it's the quiet mechanism underneath the thing so many founders describe to me – we keep shipping, and nothing moves.
The trap doesn't start with a bad decision
It starts with a good one. A customer asks for a feature. Sales says it'll help close a deal, and they're not necessarily wrong – it might. A PM writes the spec, engineering builds it, everyone celebrates the launch. Nobody is being negligent at any step in that chain. The problem is that the last step – did it actually move anything – never happens, because "launched" already got the applause and the next thing is already waiting. Multiply that by fifty requests a quarter and you have a feature factory: a busy team, a growing backlog, a great-looking velocity chart, and a product getting wider without getting better. Everyone is working hard. The revenue is flat. And because every individual decision along the way looked reasonable, nobody can point to the moment it went wrong – there wasn't one. The mechanism is structural, which is exactly why working harder inside it doesn't get you out.
"Done" means shipped, and that's most of the problem
Walk into a team that ships constantly and moves no metrics, and you'll usually find the defect in one word: done. In most teams, "done" means the feature is deployed. Code merged, release notes written, demo given, on to the next thing. The connection between that feature and the outcome it was supposed to change existed as a hypothesis in a planning doc, and it was never revisited.
The numbers on this are not subtle. One analysis of feature factories found teams ship thirty to fifty features a year but move fewer than a quarter of them to any measurable business metric. A widely-cited Pendo figure puts roughly 80% of features in the rarely-or-never-used bucket. And the AI era has made it worse, not better: a Banyan Software survey of software operators found 51% saw fewer than one in four customers use the AI features they shipped. When building gets cheaper and "done" still means "shipped," you don't get more value. You get more unused product, faster.
There's a good essay on this that reframes the whole thing as the output trap: a team can be one hundred percent on-time, one hundred percent on-budget, and zero percent on-strategy, and the dashboard will stay green the whole way. The fix it lands on is the right one – "done" has to mean shipped and measured against the outcome it was meant to move. But I'd go one step underneath that, because the reason teams don't close the loop isn't laziness. It's a belief about what a metric is for.
A metric is a hypothesis, not a target
Here's the reframe I've settled on after watching this play out on team after team: I treat every measure as the test of a hypothesis, never as a target to hit.
That sounds like a small distinction. It changes everything downstream. If the metric is a target, a flat number is a failure – something to lean on people about, to explain away, to grind harder against. If the metric is a hypothesis test, a flat number is the single most useful thing it can hand you: evidence that your picture of the world was wrong. The feature you shipped assumed a causal story – do this, and that number moves. The number came back flat. That's not the team failing. That's reality correcting your model, for free, which is exactly what you built the measurement for.
The order matters, and most teams run it backwards. They pick a metric, set it as the goal, and then build toward the number. The order that actually works: start from your read of what creates value for this customer, form a hypothesis about it, decide what you're trying to achieve, and then choose the signals that would tell you whether you were right. The metric comes last, and it comes as a question, not a quota. Because whatever you measure is what your team will read as your real priority – people read the metric, not the mission statement – and if you don't decide what counts as success on purpose, the decision gets made by default, by whichever number was easiest to track.
What changes when the metric is a test
Redefine "done" and the whole system tilts. Done stops meaning "we shipped it" and starts meaning "the metric moved – or we learned it won't, and that's information the next decision needs." A feature that shipped and didn't move its number isn't a wasted quarter; it's a disproven hypothesis, and a team that keeps its hypotheses honest gets sharper every cycle. A team that never closes the loop keeps shipping things that don't move metrics and can never quite say why.
You can see the whole shift in how the goal gets written. "Launch the new onboarding flow by March 15" is done the moment it ships – it's a feature with a date attached, and the date is the only thing it tests. "Cut new-user time-to-activation from seven days to three by the end of the quarter" isn't done until the number moves; if the new flow doesn't move it, you keep working, because the goal was the outcome and the flow was only a guess about how to get there. Same work, opposite definition of finished. One version trains a team to celebrate deploys. The other trains it to care whether the deploy mattered – and across a year, those two habits build very different products.
One guardrail, because this cuts both ways. The moment a number gets turned into a weapon – a leaderboard, a stack ranking, a stick – it has stopped doing this job, and your people will start gaming it the way any target gets gamed. Part of what you owe a team is to stand between them and that misuse. The measure exists to test your thinking, not to hunt your people. Keep it a question, and it tells you the truth. Turn it into a target, and it will tell you exactly what you demanded to hear, right up until the users are gone.
More from Consulting Operations

Consultant or Fractional CPO? You're Asking the Wrong Question
Founders compare a fractional CPO and a consultant on price and hours – but the real question is whether your gap is a decision or an operating system.

The Most Valuable Thing Research Can Tell You Is "No"
Before you build for a new market, the research that pays for itself is usually the one that returns a defensible no, not a green light.

KPMG Published a Report on Agentic AI. 40 of Its 45 Citations Were Fabricated.
I reported a bounce rate that was real and wrong, which is the small version of what just happened to four of the world's largest advisory firms.
Want help running a sharper practice?
The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.
See how this works for advisors