The Five Whys Finds About 3% of Your Problem
Researchers drew the full causal tree for one incident and found 75+ causes. A Five Whys pass finds one of them. Here's when that's fine, and when it isn't.

Some researchers took a single adverse incident, one event, and drew the complete causal tree for it – every cause and contributing factor they could establish, laid out as a structure rather than a line. It came to more than 75 of them.
Then they ran a Five Whys analysis on the same incident, the way it's normally taught. It produced one root cause. Done thoroughly, with two parallel chains instead of one, it produced two. The paper puts the yield at under 3% of the improvement opportunities the tree had identified, and works through two different Five Whys passes on the same event that arrive at two different, equally defensible answers.
I use the Five Whys. I'm not building to a conclusion that it's worthless. But that number changed how I use it, and it's worth understanding what you're getting.
Why it's such an appealing tool
The appeal is real and it isn't only laziness. Asking why repeatedly is a genuinely good instinct, it takes no training to run, it produces a legible artifact in twenty minutes, and it gives a group a shared story about what happened. That last one matters more than people admit – after a bad week, a team needs an account of events it can agree on.
It also works. For a bounded technical problem with a mechanical chain behind it, the Five Whys will frequently walk you straight to the thing. A metric moved wrong, a job failed, a feature behaves badly on one path. I reach for it in those situations and it earns its keep.
Where that 3% comes from, and what it doesn't prove
The study is from patient safety, published in BMJ Quality & Safety, and the incident is a medication error involving a wristband printer. One incident, one tree, one domain. So the honest version of the number is "in this worked example," not "Five Whys recovers 3% of causes as a general law." I'd be doing exactly what I'm criticizing if I took a single case and promoted it to a constant.
Two things make it transfer anyway. The first is that the mechanism is structural rather than clinical – a single pathway through a branching causal structure captures a small fraction of it, and that's arithmetic, not medicine. The second is that the paper's authors are careful to say healthcare is more complex than the manufacturing context the tool came from, which is the actual argument. Five Whys was built at Toyota for a production line, where a linear chain is often a fair model of reality. Ohno wasn't wrong. The tool travelled.
Software sits somewhere between a production line and a hospital, closer to the hospital than most of us like to admit, and further from it than the paper's exact number would suggest. Which is why I take the finding as directional and the reasoning as sound, and don't quote 3% as though it were measured on my team.
The two assumptions nobody states
Donald Reinertsen names them cleanly in an essay on what he calls the Cult of the Root Cause: the method assumes causality runs as a linear chain, and it assumes the beginning of that chain is the best place to intervene.
Both are sometimes true. Neither is reliably true.
The linearity assumption fails whenever failure requires several things to be wrong at once, which in software is most of the time. His fire example is the clearest version – heat, fuel, and oxygen are all necessary and none is the root cause, and fixating on one of them hides the fact that you have three possible interventions and should pick by cost rather than by position in a chain.
The intervention assumption fails more often than people expect. Reinertsen's example is seeing smoke come out of a computer: you cut the power, which treats the symptom, because response time matters more than root-cause elegance. Sometimes the cheapest durable fix is deliberately mid-chain. A spell checker is a better answer than teaching every employee to spell.
Two teams, same incident, different root cause
This is the finding that should bother you most, and it's the one least discussed.
The method isn't reproducible. Because you choose which "why" to follow at every step, and the choice is subjective, two competent teams working the same incident routinely arrive at two different root causes – and both can defend theirs, because both chains are valid. John Allspaw's version of this is the sharpest I've read: we aren't finding causes, we're constructing them. We've "found" a root cause at the moment we stop looking.
Which means the output is partly a record of who was in the room and what they already believed. A team that suspects the deploy process will find a deploy root cause. A team under pressure to show accountability will find a human one, and "human error" is the terminus of an enormous number of Five Whys chains for reasons that have nothing to do with where the fault actually sat.
The engineering write-ups from practitioners in incident response converge on the same point from the operations side: serious failures need multiple faults, each necessary and none sufficient alone. Pick one and call it the root and you've made a choice, not a discovery.
There's a second test the method doesn't have
Here's my own addition, and I'll flag it as mine rather than something the literature says: finding a true root cause is only half the job. It still has to pass a worth-solving test.
A root cause can be entirely real and still be the wrong thing to fix, if it recurs rarely, costs little when it does, and is expensive to eliminate at the source while a mid-chain mitigation is cheap. The Five Whys has no step that asks any of that. It has no step that asks whether the thing you found is worth the money. It's built to locate, not to triage, and the implicit promise that the deepest cause is the most valuable one is exactly the assumption Reinertsen is arguing with.
So two questions are worth asking after any chain terminates. What does this cost us if we never fix it, and what's the cheapest place in this chain to intervene. Sometimes those point at the root. Often they don't.
When it's still the right tool
Reach for it when the system is mechanical and bounded, when one causal chain is plausible on its face, when you need a fast shared account of a discrete failure, and when the stakes are low enough that being partly right is fine.
Be much more careful when people are in the system. A struggling team isn't a chain of five dominoes, and a root-cause pass on human performance tends to terminate on a person – which is both the least useful place to land and the one your instincts are already biased toward, for reasons I've written about when a team blamed an engineer whose tickets took 90 days.
What to do instead
For a team, I run a differential rather than a chain – several independent reads taken separately, then compared, the way a clinician generates candidate diagnoses before ruling any of them out. I use five specific lenses for that, and the point of running all of them is that each one is wrong in a different direction.
For an incident, the cheap upgrade is to stop at the first why and branch instead of descending. List everything that had to be true for this to happen, and only then pick which branches are worth walking down. You'll get a tree instead of a line, it takes maybe twice as long, and you'll be choosing between interventions rather than inheriting the one your chain happened to end on.
The Five Whys asks what caused this. That's a narrower question than what would have to change, and most of the time the second one is what you actually came to the meeting for.
More from Consulting Operations

I Wrote Down Eight Risks. All Eight Happened. Nobody Moved.
I logged eight risks against a product at Adaptavist. Within a year all eight had happened and we killed it. Writing them down changed nothing.

47% of Boutique Firms Lost a Quarter to Half Their Pipeline to ‘No Decision’
Nearly half of boutique firms lost a quarter to half their closed-lost pipeline to no decision. That isn't a closing problem. It's an evidence one.

The Five Lenses I Run Before I Diagnose a Team
A team blamed an engineer for 90-day tickets. His actual time in queue was five. Here are the five lenses I run before I diagnose a team.
Want help running a sharper practice?
The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.
See how this works for advisors