The research readout looks decisive. Twenty-four customer interviews ran in a week. Eighteen participants called the concept useful. The summary has themes, quotations, and a tidy list of requested features.
Then someone asks the question the deck cannot answer: What will these customers do differently if we build it?
Nobody knows. The interviews collected reactions to an idea, not evidence about behavior. AI made the conversations cheaper to conduct and easier to summarize, but it also made weak evidence arrive in impressive volume.
This is becoming a practical concern as AI moves from transcribing interviews to conducting them. The shift is promising. A 2026 CESifo working paper found that AI-led interviews produced thematically richer responses than other scalable survey methods, largely because the system could ask dynamic follow-up questions. The researchers also found that responses predicted behavior measured six months later.
That is more interesting than the usual claim that automation saves time. It suggests AI can help qualitative questioning reach people and situations that conventional interviews miss.
But scale changes the bottleneck. When a team can collect hundreds of conversational responses, the scarce skill is no longer arranging interviews. It is deciding what deserves to influence the product.
Fluency is not the same as depth
A good interview does not merely keep someone talking. It helps them move from a socially easy answer toward a specific account of their world.
That distinction matters because conversational AI is very good at producing the feeling of momentum. It can acknowledge an answer, paraphrase it neatly, and ask another relevant question without an awkward pause. The participant feels heard. The transcript grows. The research team receives abundant material.
Abundance can hide a thin conversation.
In a 2026 study of an AI chatbot interviewing 74 social scientists, researchers found that the bot generally stayed on topic, adapted to expertise, and encouraged detailed engagement. They also observed cases where it overinterpreted responses, asked leading questions, persisted with an ill-fitting guide, or moved too quickly. Their conclusion was admirably plain: data quantity should not define a successful qualitative interview. The team did not believe the conversations reached the depth and nuance they wanted.
The point is not that a human interviewer always goes deeper. Humans also lead witnesses, cling to scripts, reward agreeable answers, and miss the interesting sentence while planning the next question. The point is that a smooth exchange gives neither kind of interviewer proof that it found something true or consequential.
You need a way to grade the evidence inside the answer.
Use an evidence ladder
Treat each promising customer statement as the first rung of a ladder, not the finding itself. The interviewer’s job is to climb only as far as the conversation honestly allows.
1. Opinion
“That would be useful.” “I like it.” “I would probably try it.”
Opinions tell you how a concept lands in the moment. They can reveal confusing language or an immediate emotional reaction. They are weak predictors of what someone will do after the call.
2. Episode
“Last Tuesday, I copied the call notes into our CRM and then messaged the account owner because the renewal risk was missing.”
An episode anchors the conversation in something that happened. Ask when it occurred, what triggered it, who was involved, and what the person did next. Concrete sequence is usually more useful than confident generalization.
3. Constraint
“I cannot send the summary until legal reviews any pricing language, so the account owner often waits a day.”
Constraints explain why an apparently clumsy workflow survives. They include permissions, incentives, habits, dependencies, trust, and risk. A product that ignores the constraint may improve the wrong step.
4. Tradeoff
“I would accept an automated first draft, but not if it sends before I can remove negotiation details.”
Tradeoffs reveal the boundary of value. Ask what speed, control, accuracy, privacy, or effort the customer is willing to exchange. A feature is rarely simply wanted or unwanted. It becomes attractive under certain conditions and unacceptable under others.
5. Commitment
“I will bring two real call summaries to a test next week, and my manager has agreed that we can trial the workflow.”
A commitment puts some scarce resource at stake: time, access, reputation, data, budget, or a change in routine. It does not have to mean a purchase. Even a small, voluntary next step is stronger evidence than enthusiasm without cost.

The ladder is not a cross-examination. Do not force every participant toward a commitment or treat hesitation as failure. Some interviews are exploratory. Some people are excellent observers but have no authority to run a trial. The ladder simply keeps the team honest about what kind of evidence it collected.
Ask questions that move, not questions that decorate
Suppose a participant says, “I lose track of action items after customer calls.” A decorative follow-up asks, “Would automatic action items help?” The likely answer is yes, and the interview has learned almost nothing.
A moving question asks, “Tell me about the last time one got lost.” From there:
- “When did you first notice?” finds the failure point.
- “What did you do to recover?” reveals the current workaround.
- “Why not put it directly in the tracker?” exposes a constraint.
- “Which details would you never want transferred automatically?” identifies a boundary.
- “What would need to be true for you to test this on a real call?” explores commitment.
AI can be genuinely useful here. Suggested questions can remind an interviewer to pursue a missing constraint or return to an unresolved detail while the conversation is still live. The human still decides whether the moment calls for another probe, a pause, a change of subject, or simple empathy.
That division of labor protects what is valuable about real-time assistance. The copilot can track threads and candidate questions. The interviewer can watch the participant, sense discomfort, and decide which relationship matters more than completing the guide.
Do not turn themes into votes
After many interviews, AI can group similar answers quickly. That is helpful for navigation, but a frequent theme is not automatically an important one.
Ten participants may praise a dashboard they have never used. One may describe a compliance rule that makes the entire concept impossible for the target market. Counting mentions would favor the praise. Product judgment should favor the constraint.
For each proposed finding, keep four things together:
- Claim: what the team currently believes.
- Strongest evidence: the clearest observed episode, constraint, tradeoff, or commitment.
- Counterexample: the best evidence that the claim may not hold.
- Confidence: what is known, what is inferred, and what should be tested next.
This small record also makes summaries safer. A polished synthesis can otherwise erase the path from raw conversation to product conclusion.
There is a related warning for synthetic research. In a 2026 product-discovery study, researchers built generative agents from interviews with 51 knowledge workers and compared the agents’ reactions to new concepts with the humans’ own responses. The agents approximated response patterns at the population level but were imprecise representations of specific individuals. That makes simulation potentially useful for early screening, not a substitute for hearing what a particular customer notices, fears, or chooses.
Put scale in the right place
Use AI-moderated interviews when the questions are reasonably understood, broader participation matters, and consistent coverage is valuable. Keep a person close to messy discovery, sensitive topics, high-stakes decisions, and moments where an unexpected answer should overturn the guide.
For human-led conversations, a meeting copilot can create a useful middle path. It can preserve a live recap, hold context, and suggest questions without replacing the person responsible for trust and interpretation. Afterward, persistent memory can help the team compare episodes and constraints across calls, provided people still inspect the evidence behind each theme.
That is the useful role for a tool like Caspi: help the interviewer remember more and notice more while leaving the judgment where it belongs. More interviews can expand what a team hears. The evidence ladder decides what the team has actually learned.