Customer Discovery & the Mom Test
- Explain why polite people lie in discovery and how question design defeats it
- Convert hypothetical and opinion questions into past-behavior and specifics questions
- Extract workflow, pain frequency, cost, and current workarounds from an interview transcript
- Map stakeholders β economic buyer, users, champion, blocker β from conversational evidence
- Elicit testable success criteria ("how would you know this worked?") without leading the witness
| Spaced-rep warm-up: due cards + Day 161 self-quiz remediation check | 10 min |
| ELI5 + tech read; write the three rules and five slots from memory | 15 min |
| Guided: Harbor & Vane annotation + Nortia surgery | 40 min |
| Practice: live discovery with the simulated customer | 20 min |
| Project: assemble the discovery kit | 25 min |
| Quiz + flashcards | 10 min |
Builds on: Day 162 β The FDE role β the consultant hat Β· Day 105 β Talking to non-technical stakeholders
Ask your mom "would you use an app that plans family dinners?" and she'll say "of course, sweetheart!" β because she loves you, not because she'd use it. Ask instead "how did you decide what to cook last Tuesday?" and she cannot lie: either a story about a real struggle tumbles out, or a shrug reveals there is no problem here at all. That is the Mom Test, and the insight generalizes brutally: everyone you interview is polite, and questions about the future or about opinions ("would youβ¦", "do you thinkβ¦") harvest politeness. Questions about the past and about specifics ("walk me through the last timeβ¦", "what did that cost you?", "what did you try instead?") harvest facts.
For an FDE the stakes are a hundred times a dinner app. A VP says "we need AI for our claims process" β and if you ask "would a summarization assistant help?" the answer is yes, it is ALWAYS yes, and three months later you demo a tool that nobody opens. The discovery interview is the cheapest, highest-leverage engineering you will ever do: an hour of well-designed questions routinely saves a quarter of well-engineered waste.
The most expensive failure in applied AI is not hallucination β it is the perfectly engineered solution to the wrong problem. Enterprises are littered with abandoned AI pilots that demoed beautifully and solved nothing anyone was paid to care about. FDE interview loops test discovery directly: you will role-play a customer conversation and be graded on question quality (Day 168 simulates this; the Exponent guide's case round is exactly this). And the skill compounds: every requirements doc (Day 164), proposal (165), and demo (167) inherits its truth or its fiction from the discovery that fed it.
Guided practice
Annotate a real-shaped transcript: Harbor & Vane Insurance
25 minBelow is an excerpt from a discovery call with Dana Okafor, Claims Operations Manager at Harbor & Vane, a 400-person commercial-insurance brokerage. Read once for flow. Then annotate line by line: mark each interviewer question MT-GOOD (past/specific/their-life) or MT-BAD (hypothetical/opinion/pitch), and mark each Dana answer for which extraction slot it fills (workflow / pain-quant / workaround / stakeholder / success).
FDE: Thanks for the time, Dana. Before anything else β could you walk me through the last claim that came in yesterday or today? Just what literally happened, start to finish. DANA: Sure. A water-damage claim from a logistics client came in around 9am β email to our claims inbox, PDF attachments, maybe 30 pages. Priya on my team opened it, skimmed for the policy number, looked it up in AS400 β that's our policy system β then started keying the loss details into ClaimCore, our claims platform. That took her until, I want to say, 11:30. FDE: What was she actually doing for those two and a half hours? DANA: Reading, mostly. The loss runs are never in the same format. Adjusters' reports, photos, invoices β she has to find dates, amounts, cause-of-loss language, and policy exclusions that might apply, and re-type them. And she double-checks against the policy PDF because AS400's data is... let's say vintage. FDE: How many claims like that land in a week? DANA: Complex commercial ones? Forty to sixty across the team. Simple auto stuff is maybe two hundred but those are quick. FDE: You said Priya double-checks against the policy PDF β tell me about the last time skipping that check caused a problem. DANA: [laughs] March. We paid a claim that an exclusion clearly barred β nobody caught the endorsement. That was a $40K mistake and I spent a week of my life on the postmortem. FDE: What have you tried already, for the re-keying part? DANA: We piloted an OCR tool two years ago. It choked on the adjusters' scanned reports and the team quietly went back to doing it by hand. IT still brings it up β they bought licenses. Oh, and don't repeat that. FDE: Noted. If the intake time went from two and a half hours to thirty minutes, what number changes, and who sees it? DANA: Cycle time to first response. It's on my dashboard and my VP's β we promise brokers 48 hours and we hit maybe 70%. If that went to 90 I'd... honestly I'd take the win to Marcus myself. He owns the ops budget. FDE: Would an AI copilot that pre-extracts the loss details into ClaimCore be something the team would use? DANA: Oh β sure, that sounds great.
- Annotate every FDE turn. (You should find exactly one MT-BAD question β the final one. Note what the answer is worth: nothing. Note also how much the "what have you tried" question yielded: a failed OCR pilot, a wary IT department holding budget scar tissue, and a confidentiality boundary.)
- Fill the five extraction slots from evidence only; quote the supporting line for each.
- Stakeholder map: Dana (champion? user?), Priya (user), Marcus (economic buyer), IT (blocker with a grudge β name why). What is your evidence for each label?
- Write the one question you would ask next, and defend it against the three rules.
Question surgery on a corrupted interview: Nortia Labs
15 minThis excerpt is from a DIFFERENT discovery call β a biotech scale-up, Nortia Labs β conducted badly. Rewrite it.
FDE: We've built an incredible AI platform for lab documentation. I think it could transform your workflows. Do you feel your scientists spend too much time on documentation? DR. IKEDA: I suppose everyone would say yes to that. FDE: If we could cut documentation time in half, would that be valuable to Nortia? DR. IKEDA: In principle, of course. FDE: Would your scientists use an AI assistant that drafts their experiment write-ups automatically? DR. IKEDA: Maybe. Some might. We're quite busy with the FDA submission this quarter, though. FDE: Great β so would next week work for a demo? DR. IKEDA: ...Send me some materials and we'll see.
- Label the failure in each FDE turn (pitch-first, hypothetical, leading, compliment-fishing, premature advance).
- Note the two genuine signals the FDE talked past: "busy with the FDA submission" (their ACTUAL top pain this quarter β a better wedge than documentation) and the soft brush-off ending (no time, no introduction, no commitment = polite no).
- Rewrite the interview as five questions that obey the rules, starting from the FDA-submission thread. For each, say which extraction slot it targets.
- One paragraph: why "send me some materials" is a failure outcome, and what commitment-currency a good version would have asked for instead.
On your own
Live discovery against a simulated customer
20 minRun a 15-minute discovery interview with your AI tutor playing a role you must not peek at: "You are Sam Torres, Head of Customer Support at Fieldstone Software (B2B payroll, 900 customers, 12 support agents). You have a real, specific operational pain involving your support workflow that you will only reveal through concrete stories when asked good questions. Politely deflect hypotheticals with vague agreement. Volunteer nothing. If asked Mom-Test-style past-behavior questions, answer with rich specifics including numbers, system names, and people."
Constraints: no pitching, no solution words ("AI", "copilot", "tool") for the first ten minutes; fill all five extraction slots; end by asking for ONE commitment and note what you get.
Afterward, self-grade: count MT-GOOD vs MT-BAD questions from your own transcript; check which slots are actually filled with QUOTED evidence rather than your inference. Hints if stuck mid-interview: "walk me through yesterday's worst ticket", "what happened next?", "what does that cost you in a week?", "what have you tried?"
The discovery kit
Build docs/discovery_kit.md β the artifact you will carry into Day 168's simulation and real interviews: (1) your annotated Harbor & Vane transcript with the five slots filled and quoted; (2) the Nortia rewrite (five repaired questions, slot-tagged); (3) a personal question bank of 15 questions organized by extraction slot, each obeying the three rules β at least 5 written by you, not copied from today's text; (4) the stakeholder-map template (role, name, pain exposure, power, evidence line); (5) your self-graded score from the live practice interview with two specific improvements for next time.
Common mistakes & misconceptions
- Pitching in the first five minutes. Once your idea is on the table, every subsequent answer measures politeness toward you, not reality.
- Accepting compliments as validation. "That sounds great" (see Dana's last answer) is a zero-information response to a bad question; only past behavior and commitments count.
- Asking about the future: "would you / do you think / how likely." Humans are terrible predictors and generous liars about their future selves; convert to "last time" questions.
- Ignoring the workaround question. What they currently do (including a failed OCR pilot!) reveals pain intensity, budget history, and political scar tissue in one answer.
- Interviewing only your champion. Dana is not the buyer (Marcus) nor the daily user (Priya) nor the blocker (IT) β a deal map built from one enthusiastic voice collapses at security review.
- Ending without asking for a commitment. Time, an introduction, data access β if the meeting ends in compliments and "send materials," you learned nothing about whether they mean it.
Q1. Which question yields the most trustworthy discovery data?
Q2. In the Harbor & Vane call, Dana's "Oh β sure, that sounds great" in response to the copilot question is best treated asβ¦
Q3. A prospect ends discovery with warm compliments and "send me some materials." Under commitment-currency rules this isβ¦
Go deeper β curated resources
- Interview a real person tonight β Run a 15-minute Mom-Test interview with a friend about any workflow of theirs (expense reports, meal planning). The rules feel obvious on paper and evaporate under social pressure; only live reps fix that.
- Both transcripts fully annotated with slot-quoted evidence
- Live interview run; MT-GOOD/BAD self-count recorded
- discovery_kit.md committed with a 15-question original-heavy bank
- Quiz β₯ 2/3
β Back: This is Day 162's consultant hat given a method, and it rhymes with Day 134's eval mindset: in both, you design the measurement BEFORE trusting the output β questions are to customers what golden sets are to models.
Forward β: Tomorrow (Day 164) the Harbor & Vane transcript becomes a requirements doc. Day 168's simulation opens with a fresh transcript to mine, and Day 174 adds live objections. The success-criteria question returns as pilot metrics on Day 173.
Unlocks: D164 Ambiguity β Requirements Β· D168 Week 24 Checkpoint: FDE Simulation I