AI can help product discovery.
I use it. I think teams should use it. It can clean up notes, compare interviews, challenge a hypothesis, and turn a pile of messy evidence into something another human can actually read.
But AI cannot do product discovery for you.
That distinction sounds obvious until a team has a polished persona, a jobs-to-be-done document, a market summary, a feature map, and a roadmap before speaking to a single user.
Nothing in those documents has to look ridiculous.
That is the problem.
The output is plausible enough to feel like progress and complete enough to remove the discomfort that should have forced the team into reality.
The Productivity Trap
Imagine a team has an idea for a new workflow.
They open an LLM and ask for:
- target personas;
- user pain points;
- jobs to be done;
- likely objections;
- feature opportunities;
- interview questions;
- a prioritized roadmap.
An hour later, the team has a clean document.
The personas have names. The pain points sound familiar. The roadmap has sensible phases. The language is professional. Stakeholders can react to it immediately.
It feels much better than uncertainty.
But the document contains a hidden category error.
The team asked a model to describe what users are like and then began treating the answer as evidence of what users are like.
A plausible answer is not the same as a discovered truth.
Evidence, Inference, and Invention
A useful way to protect discovery work is to separate three things that often get blended together.
Evidence
Something observed in reality.
A user said it. A user did it. Product data shows it. Support tickets repeat it. A domain expert demonstrated the workflow. A customer abandoned the process at the same point three times.
Evidence can still be incomplete or misleading, but it has contact with reality.
Inference
What the team believes the evidence may mean.
A user repeatedly exports data before performing another task. The team infers that the current workflow does not provide enough confidence or control.
Inference is unavoidable. Product work requires interpretation.
The important part is that an inference should remain connected to the evidence that produced it.
Invention
A gap filled because the prompt requested a complete answer.
The model creates a persona, assigns motivations, predicts objections, or explains a workflow without direct evidence.
Invention is not always useless. It can help a team imagine possibilities.
It becomes dangerous when it is formatted like evidence and stored next to evidence without a label.
The model may not know which parts came from transcripts, which came from general patterns, and which it created to make the document coherent.
The team needs to know.
AI Should Work Downstream From Evidence
My rule is simple:
AI should process evidence, not manufacture contact with reality.
That means the model becomes more useful after humans have done at least some real discovery.
Give it interview transcripts, support conversations, product events, domain notes, and competing interpretations. Ask it to organize, compare, challenge, or identify gaps.
Do not give it an empty page and ask it who your users are.
The difference is not philosophical. It changes the quality of the decision.
When AI works downstream from evidence, its output can be traced back.
When it works upstream from evidence, the output often becomes a sophisticated guess that the team later forgets was a guess.
Where AI Genuinely Helps
There are several places where AI can make discovery work better without replacing it.
Preparing for conversations
AI can review a hypothesis and help identify leading questions.
Instead of asking:
Would automatic reconciliation save you time?
It can help reframe the conversation toward behavior:
Walk me through the last time you reconciled this kind of transaction. Where did you stop to check something manually?
It can also ask what evidence would disprove the team’s current belief.
That is useful because teams are naturally better at collecting confirmation than contradiction.
Organizing real interviews
A model can summarize a transcript, extract repeated themes, group similar frustrations, and compare language across several interviews.
This is especially useful when the raw material is large and inconsistent.
The summary should still point back to the source. A theme without quotes, timestamps, or interview references becomes hard to verify and easy to exaggerate.
Finding contradictions
Discovery is not only about repeated patterns.
Sometimes the valuable signal is that two users who appear similar behave completely differently.
AI can help surface:
- conflicting expectations;
- different meanings for the same domain term;
- exceptions to the dominant pattern;
- statements that do not match observed behavior;
- unanswered questions shared by several interviews.
The contradiction is often where the product problem becomes interesting.
Stress-testing an interpretation
After the team forms a conclusion, AI can argue against it.
Not with a vague “challenge this,” but with evidence-aware prompts:
- Which observations do not support this conclusion?
- What alternative explanation fits the same evidence?
- Which user segment might experience this differently?
- What would we need to verify before turning this into a requirement?
This does not create truth. It reduces the chance that the first coherent interpretation becomes the only interpretation.
Improving communication
Discovery material is usually messy.
AI can turn notes into a concise decision record, summarize the evidence for stakeholders, or produce several levels of detail for different audiences.
That is acceleration.
It becomes replacement only when the polished summary is allowed to hide uncertainty.
Where Teams Lose Contact With Reality
The failure modes are subtle because each one can look like good product work.
Synthetic personas
A persona generated from market language and broad demographic assumptions may look realistic.
But nobody has met that person.
The persona can be used as a hypothesis generator. It should not become a substitute customer whose invented needs settle product arguments.
Confirmation as a service
A team asks whether an idea makes sense, and the model produces a thoughtful explanation of why it does.
The answer may include risks, but it usually accepts the framing of the prompt.
The team feels challenged because the response is nuanced. In reality, the model may only be helping them create a stronger version of their existing belief.
Premature synthesis
The team has two interviews and asks for the main patterns.
The model produces five patterns because the requested format implies that patterns must exist.
A human might have looked at the same material and said, “We do not know yet.”
That sentence is less impressive and often more valuable.
Artifact theatre
The discovery deck looks finished, so discovery feels finished.
There are diagrams, themes, opportunity areas, confidence scores, and a roadmap.
The visual completeness of the artifact gets confused with the completeness of the work.
Roadmap laundering
An unsupported assumption enters an AI-generated document. The document becomes stakeholder alignment. Alignment becomes tickets. Tickets become a roadmap.
By the time engineering sees the requirement, the original guess has passed through enough professional formats to look earned.
That is how teams build around features and expectations that do not actually exist.
A Practical AI-Assisted Discovery Workflow
The goal is not to keep AI outside discovery.
The goal is to keep evidence in control of it.
Before speaking to users
Use AI to:
- list the assumptions inside the product idea;
- identify which assumptions are riskiest;
- draft neutral interview questions;
- flag leading or solution-biased language;
- generate alternative explanations;
- define what evidence would disprove the hypothesis.
Do not use it to decide what users believe.
During discovery
Keep the conversation human.
Listen for hesitation, confusion, workarounds, contradictions, and the moments where somebody reaches for another tool.
A live transcript can be useful. A live summary should not replace attention.
The strange pause before an answer may contain more information than the answer itself.
After each conversation
Record:
- what was observed;
- what was said directly;
- what surprised the team;
- what appears to be an inference;
- what needs a follow-up;
- which previous belief became weaker.
Then use AI to structure the material.
After several conversations
Ask AI to compare the evidence, but require source references.
A useful output is not only “these are the main themes.”
It should also include:
- supporting examples;
- contradicting examples;
- confidence level;
- missing segments;
- questions the evidence cannot answer;
- proposed next validation step.
Before turning discovery into delivery
Every important claim should pass a reality check:
- What evidence supports it?
- Where can that evidence be inspected?
- Is this a direct observation or an inference?
- Did AI introduce any part of the claim?
- What contradicts it?
- What still needs validation?
- What is the cost of being wrong?
If the claim cannot survive those questions, it is not ready to become a confident requirement.
The Red-Team Record
A small template can stop a generated interpretation from silently becoming product truth.
Claim
What are we currently assuming?
Supporting evidence
- Interview
- Observed behavior
- Product data
- Support request
- Domain input
Contradicting evidence
What does not fit the claim?
Interpretation
What do we think the evidence means?
AI contribution
Did AI summarize, infer, challenge, or generate any part of this claim?
Confidence
Low / Medium / High
Required validation
What must happen before this becomes a product decision?
The point is not bureaucracy.
The point is keeping the chain visible.
When the team later asks why a feature exists, it should be able to find the evidence, the interpretation, and the decision without reconstructing the story from memory.
Discovery Is Not a Document
Discovery is often discussed as a phase that produces artifacts.
The artifacts matter. They help teams remember, communicate, and decide.
But discovery itself is contact with reality.
It is the conversation that exposes a false assumption. The workflow demonstration that makes the mockup look naive. The support ticket that shows a supposedly rare edge case happens every day. The domain expert who says the technically elegant solution breaks the actual business process.
AI can summarize all of that.
It cannot be the source of all of that.
This is the same boundary I was trying to describe in Up Is Not Up When You’re Stuck in an Avalanche: speed helps until it removes the friction that kept the team close to truth.
Closing
AI is useful in product discovery when it helps a team see evidence more clearly.
It is dangerous when it helps a team stop noticing that evidence is missing.
Use it to prepare better questions.
Use it to organize messy conversations.
Use it to find contradictions and challenge the first interpretation.
Use it to make the work legible.
But do not ask it to invent the users, interview them, summarize them, and validate the roadmap in the same loop.
That is not accelerated discovery.
That is a team becoming more efficient at agreeing with itself.