If Your Best Customer and Your Worst Prospect Give the Same Answer, the Question Tested Nothing
A validation question earns its place only if your best customer and your worst-fit prospect would answer it differently. Most don't.

Here's a question that shows up in a lot of customer interview scripts: "Would you use a tool that automatically flagged which of your accounts were at risk of churning?"
Ask it of your best customer, the one who renews every year and sends you referrals, and you'll probably get a yes. Ask it of a prospect who's completely wrong for your product – wrong size, wrong budget, no real churn problem – and you'll probably get a yes there too. People are polite, the tool sounds useful, and saying yes to a hypothetical costs nothing.
When both of those people give you the same answer, the question didn't test anything. It produced a data point that feels like validation and can't separate the people who'll pay from the people who won't, which is the only thing a validation question is for.
The clearest version of this idea I know comes from a job that had nothing to do with customer discovery.
What six certification exams taught me about questions that test nothing
Between 2015 and 2017, while I was at Adaptavist, I helped design and launch six Atlassian Professional Certifications. These weren't multiple-choice trivia quizzes. They were meant to test whether someone could actually administer Jira, manage Confluence or design workflows in production, and you needed a couple of years of real hands-on experience to pass the administration exams.
I'm not a psychometrician. We worked with one from Alpine Testing, who helped build each exam's blueprint and ran the beta tests, then fed back how every question and answer had performed. My job was to revise the questions based on that data.
The beta process used item analysis on every question: a difficulty index, a discrimination index and an analysis of the wrong answers. The principle underneath all of it is simple. A question that experts and novices both get right isn't testing anything useful, because it can't tell you who knows the material. So questions were cut aggressively based on the data.
Writing questions that pass that bar is harder than it sounds. A badly written question ends up testing reading comprehension instead of knowledge. A vague one lets an expert and a novice arrive at the "right" answer for completely different reasons. And in Jira specifically, you have the added problem that the tool lets you do almost anything multiple ways, so it's easy to write a question with more than one correct answer without noticing.
The discrimination test
In assessment design, a question's discrimination index compares how the high scorers and the low scorers on the whole test did on that one item. Florida State University's testing office describes it as a number between −1 and 1, where 0.3 and above is good and values close to zero "mean that most students performed the same on an item." Penn State's guide to item analysis makes the point that matters most here: very easy questions are poor discriminators, because "when most students get the answer correct… it is difficult to ascertain who really knows the content."
A hypothetical discovery question is a very easy question. Almost everyone gets it "right," in the sense that almost everyone says yes.
The analogy isn't a perfect match. On an exam, the high and low groups come from people's scores on the same test, and the index is a real statistic computed across a large beta cohort. In discovery you're having a handful of conversations, and the groups are defined by fit – customers who pay and stay versus prospects who are wrong for you – which you decide in advance. Nobody should be computing a discrimination index on ten interviews. What carries over is the logic, and it turns into one question you can ask about any question in your script:
If my best-fit customer and my worst-fit prospect each answered this, would their answers be different?
If not, the question measured politeness. It might still be useful for building rapport or understanding context, but it can't validate a direction.
Why hypothetical questions can't pass it
Hypothetical questions fail the test for a reason the research has documented for a long time: what people say they'd do about something that doesn't exist yet is a weak guide to what they'll actually do.
Morwitz, Steckel and Gupta studied when stated purchase intentions actually predict sales, and found that intentions correlate with purchase less for new products than for existing ones, and less when people are asked about a whole product category than about a specific product. "Would you use a tool that does X?" is both of those at once – a new product, described as a category. It's the weakest form of the question you could ask.
The willingness-to-pay research points the same way. A 2020 meta-analysis in the Journal of the Academy of Marketing Science pooled 77 studies and found that people's hypothetical willingness to pay ran about 21% above what they'd really pay, with bigger gaps for higher-value and specialty products. That measures how much people inflate a price rather than whether they'd say yes at all, so I wouldn't stretch it further than the direction it shows. Hypothetical answers run high.
Put those together and you get the problem with the churn-tool question. Your best customer's yes and your worst-fit prospect's yes are both inflated, both weakly predictive, and nearly indistinguishable from each other. That's how a question that can't discriminate manufactures a cheap yes. A cheap yes leads to an expensive no, and in product work the no usually shows up after the build, when the people who said they'd use it don't.
Questions that do discriminate
The questions that pass the test are mostly about the past and about cost, because those are the things a good-fit and a bad-fit prospect really do answer differently.
Rob Fitzpatrick's The Mom Test puts the core rule plainly: ask about specifics in the past instead of generics or opinions about the future. "When was the last time an account churned on you, and what did you do in the week before it happened?" gets very different answers from your best customer and from a prospect who doesn't have the problem. One of them has a story with names and dates in it. The other one has to think about it.
Questions about what someone has already spent work the same way. "What are you using to catch this today, and what does it cost you?" Someone with the real problem has an answer, even if the answer is a spreadsheet and two hours every Monday. Someone without it shrugs.
And questions that ask for a commitment discriminate best of all. As Sachin Rekhi summarizes the book, a commitment means giving up something you value – time, reputation or money. Asking for a working session with the person who owns the problem, an introduction to their boss, or a paid pilot will sort your best-fit and worst-fit prospects almost instantly, because only one of them has a reason to say yes.
Where the test stops
Not every question in an interview has to discriminate.
The testing guides make the same point about exams. Penn State notes that easy items sometimes stay on purpose, and the University of Minnesota's item-analysis guide cautions that item statistics "should probably be used more for item improvement than for discarding items." Some interview questions exist to get someone talking, or to learn how their team is organized, or to understand their week. Those don't need to separate anyone.
The test applies to the questions you're going to lean on – the ones whose answers will end up in a slide that says "customers want this," justifying a direction bet. Those are the ones that need to earn their place, and the easiest way to check is to imagine your best customer and your worst-fit prospect answering side by side.
On the certification exams, a question everyone got right was cut because it couldn't tell an expert from a novice. Your validation questions deserve the same standard, because the people deciding what to build are relying on them to tell a customer from someone being polite.
More from Consulting Operations

31% of Expert-Network Experts Say They've Taken Calls They Weren't Qualified For
A survey of 1,368 network experts found 31% took calls they weren't qualified for. If your name is on the recommendation, the call isn't the evidence.

Every Language We Launched Grew 70–120% a Week. The Demand Was Visible Before We Built Anything.
The best market validation is often demand you can already see in what people do. Here's how we found it at CoinDesk before building anything.

Five Months Late, They Expanded the Offshore Team to 60 Engineers. It Got Later.
A late team kept adding people and kept getting slower. Brooks's law explains part of it. What it leaves out is what late teams actually lack.
Want help running a sharper practice?
The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.
See how this works for advisors