Role Playbooks4 min read

Turn Customer Interviews Into Testable Product Hypotheses With AI

A four-stage evidence ladder helps product teams use AI without confusing interview observations, interpretations, hypotheses and tests.

Bokili Editorial· Verified September 16, 2026
ShareX
Interview evidence moving through observation, interpretation and hypothesis stages towards a small product test

Customer interviews contain evidence, but an AI summary can turn that evidence into a confident product story too quickly. A sentence such as “people want more control” may hide several leaps: what the participant actually did, what the researcher inferred, what the team believes, and what should be tested next. For product managers and user researchers, the useful job is not to make the notes sound certain. It is to keep those layers visible.

Use AI as a sorting partner, not as the person who decides what the research means. The workflow below preserves source markers, labels every inference and ends with a small test. A human researcher still checks the interview record, weighs contradictory evidence and chooses whether a hypothesis deserves attention.

The evidence-to-test ladder

1

1. Observation

Record what the participant said or did, with a transcript line, timestamp or note reference. Remove interpretation.

2

2. Interpretation

State what the observation may mean. Use tentative language and keep plausible alternatives.

3

3. Hypothesis

Write a falsifiable claim about a user, situation and expected behaviour. Tie it to the observations that support it.

4

4. Test

Choose the smallest next activity that could strengthen or weaken the claim, with a signal you can observe.

Start with observations, not themes

The UK Government Service Manual recommends separating what a team saw and heard from its interpretation, then grouping observations to identify findings and actions. That order matters. If the first instruction to an AI system is “find the main themes”, it may combine different behaviours under one neat label and erase the trail back to the interview.

Instead, ask for one observation per row. Require a source marker and a short verbatim excerpt where the record allows it. Tell the system to write “missing” rather than reconstructing details. Then sample several rows against the transcript before moving on. This makes the model’s omissions and paraphrases easier to spot.

Premature conclusionCheckable research step
Interview noteUsers need spreadsheet exportA participant copied receipts into a spreadsheet before submitting an expense
MeaningThe feature request is obviousThe behaviour may indicate a need to review or control data before submission
Next moveAdd an export featureTest whether an editable review step changes completion and correction behaviour

Make every inference earn its place

A practical AI-assisted workflow

  1. 1

    Prepare a safe evidence set

    Remove personal or sensitive details you do not need. Include stable source markers so every output can be checked.

  2. 2

    Extract observations only

    Ask for actions, quotations and stated constraints. Prohibit themes, recommendations and invented context at this stage.

  3. 3

    Generate competing interpretations

    For each observation, request two or three plausible meanings and the evidence that would distinguish them.

  4. 4

    Draft falsifiable hypotheses

    Use an if–then–because shape. Include the target user and situation, then name what would count against the claim.

  5. 5

    Design the smallest test

    Choose an interview question, prototype task, log review or other low-cost activity with an observable signal and a decision point.

Reading is a start. Practice makes it stick.

Start learning

Worked example: do not confuse a workaround with a feature request

Imagine a fictional interview about a business expense product. The participant exports receipt details to a spreadsheet before final submission. An AI summary might label this “demand for spreadsheet export”. That is only one interpretation. The participant may instead be checking totals, correcting categories or creating a record they trust.

The observation is the export-and-check behaviour. One interpretation is that the participant needs control before committing the expense. A testable hypothesis becomes: “When finance coordinators can review and edit extracted receipt details before submission, they will correct errors inside the product rather than create a separate spreadsheet, because the review step provides a trusted checkpoint.”

The smallest next test could be a clickable prototype with an editable review screen. Ask several relevant users to complete the same task. Observe whether they notice and correct planted errors, whether they still create an external file and what evidence they say they need before submitting. Decide in advance what result would support, weaken or redirect the hypothesis.

Quality gate before sharing a hypothesis

  • Every observation has a source marker that a reviewer can open.
  • Interpretations are labelled as tentative, not presented as participant statements.
  • At least one credible alternative interpretation is recorded.
  • The hypothesis names a user, situation, expected behaviour and reason.
  • The test includes an observable signal and a result that could count against the claim.
  • A researcher has checked the AI output against the original evidence.
Ten-minute evidence-ladder exercise
  1. Choose one deidentified interview excerpt and mark the exact words or behaviour you observed.
  2. Write two different interpretations without choosing a winner.
  3. Turn the stronger interpretation into one falsifiable if–then–because hypothesis.
  4. Name the smallest test and one result that would weaken the hypothesis.
  5. Ask a colleague to trace each claim back to the excerpt.

Protect the evidence boundary

Do not upload raw research containing personal, confidential or sensitive information to a tool unless your organisation has approved that use. Minimise the data first and follow your research and AI policies.

Connect the workflow to the next decision

This ladder complements Bokili’s guides to turning customer complaints into evidence-first briefs, checking AI summaries for omissions and defining done before delegating work to AI. Each practice keeps a human decision attached to checkable evidence instead of polishing uncertain output.

Sources

  1. Analyse a research sessionGOV.UK Service Manual
  2. How the discovery phase worksGOV.UK Service Manual
  3. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNIST
ShareX

Reading is a start. Practice makes it stick.

Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.

Start learning

Keep reading