Turn Customer Interviews Into Testable Product Hypotheses With AI
A four-stage evidence ladder helps product teams use AI without confusing interview observations, interpretations, hypotheses and tests.

Customer interviews contain evidence, but an AI summary can turn that evidence into a confident product story too quickly. A sentence such as “people want more control” may hide several leaps: what the participant actually did, what the researcher inferred, what the team believes, and what should be tested next. For product managers and user researchers, the useful job is not to make the notes sound certain. It is to keep those layers visible.
Use AI as a sorting partner, not as the person who decides what the research means. The workflow below preserves source markers, labels every inference and ends with a small test. A human researcher still checks the interview record, weighs contradictory evidence and chooses whether a hypothesis deserves attention.
The evidence-to-test ladder
1. Observation
Record what the participant said or did, with a transcript line, timestamp or note reference. Remove interpretation.
2. Interpretation
State what the observation may mean. Use tentative language and keep plausible alternatives.
3. Hypothesis
Write a falsifiable claim about a user, situation and expected behaviour. Tie it to the observations that support it.
4. Test
Choose the smallest next activity that could strengthen or weaken the claim, with a signal you can observe.
Start with observations, not themes
The UK Government Service Manual recommends separating what a team saw and heard from its interpretation, then grouping observations to identify findings and actions. That order matters. If the first instruction to an AI system is “find the main themes”, it may combine different behaviours under one neat label and erase the trail back to the interview.
Instead, ask for one observation per row. Require a source marker and a short verbatim excerpt where the record allows it. Tell the system to write “missing” rather than reconstructing details. Then sample several rows against the transcript before moving on. This makes the model’s omissions and paraphrases easier to spot.
| Premature conclusion | Checkable research step | |
|---|---|---|
| Interview note | Users need spreadsheet export | A participant copied receipts into a spreadsheet before submitting an expense |
| Meaning | The feature request is obvious | The behaviour may indicate a need to review or control data before submission |
| Next move | Add an export feature | Test whether an editable review step changes completion and correction behaviour |
Make every inference earn its place
A practical AI-assisted workflow
- 1
Prepare a safe evidence set
Remove personal or sensitive details you do not need. Include stable source markers so every output can be checked.
- 2
Extract observations only
Ask for actions, quotations and stated constraints. Prohibit themes, recommendations and invented context at this stage.
- 3
Generate competing interpretations
For each observation, request two or three plausible meanings and the evidence that would distinguish them.
- 4
Draft falsifiable hypotheses
Use an if–then–because shape. Include the target user and situation, then name what would count against the claim.
- 5
Design the smallest test
Choose an interview question, prototype task, log review or other low-cost activity with an observable signal and a decision point.
Reading is a start. Practice makes it stick.
Start learningWorked example: do not confuse a workaround with a feature request
Imagine a fictional interview about a business expense product. The participant exports receipt details to a spreadsheet before final submission. An AI summary might label this “demand for spreadsheet export”. That is only one interpretation. The participant may instead be checking totals, correcting categories or creating a record they trust.
The observation is the export-and-check behaviour. One interpretation is that the participant needs control before committing the expense. A testable hypothesis becomes: “When finance coordinators can review and edit extracted receipt details before submission, they will correct errors inside the product rather than create a separate spreadsheet, because the review step provides a trusted checkpoint.”
The smallest next test could be a clickable prototype with an editable review screen. Ask several relevant users to complete the same task. Observe whether they notice and correct planted errors, whether they still create an external file and what evidence they say they need before submitting. Decide in advance what result would support, weaken or redirect the hypothesis.
Quality gate before sharing a hypothesis
- Every observation has a source marker that a reviewer can open.
- Interpretations are labelled as tentative, not presented as participant statements.
- At least one credible alternative interpretation is recorded.
- The hypothesis names a user, situation, expected behaviour and reason.
- The test includes an observable signal and a result that could count against the claim.
- A researcher has checked the AI output against the original evidence.
- Choose one deidentified interview excerpt and mark the exact words or behaviour you observed.
- Write two different interpretations without choosing a winner.
- Turn the stronger interpretation into one falsifiable if–then–because hypothesis.
- Name the smallest test and one result that would weaken the hypothesis.
- Ask a colleague to trace each claim back to the excerpt.
Protect the evidence boundary
Do not upload raw research containing personal, confidential or sensitive information to a tool unless your organisation has approved that use. Minimise the data first and follow your research and AI policies.
Connect the workflow to the next decision
This ladder complements Bokili’s guides to turning customer complaints into evidence-first briefs, checking AI summaries for omissions and defining done before delegating work to AI. Each practice keeps a human decision attached to checkable evidence instead of polishing uncertain output.
Sources
- Analyse a research session — GOV.UK Service Manual
- How the discovery phase works — GOV.UK Service Manual
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile — NIST
Reading is a start. Practice makes it stick.
Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.
Start learningKeep reading

Corporate AI Training Needs an Exception Library
Corporate AI training should rehearse recurring boundary cases, so employees know when to proceed, pause for evidence or escalate.

AI Literacy Training for Business: Teach Three Levels of Authority
AI literacy training for business becomes actionable when every workflow states whether AI may advise, prepare or act—and what human control each level needs.

Decide What Saved Time Becomes Before You Claim AI ROI
AI ROI does not appear when a timer stops. Turn released capacity into one owned business outcome, then verify that the outcome happened.