GuiaAI interviewsinterview guideprobe designqualitative research

How to Write an AI Interview Guide: Probing for Evidence, Not Reinforcing Assumptions

Translate research questions into topic objectives, primary questions on concrete experiences, neutral evidence probes, transitions, and stopping conditions. Pre-test each version using coverage, completeness, leading, and safety metrics.

Atualizado recentemente 2 de setembro de 2026 33 min de leitura

A usable AI interview guide is not a list of questions, nor a single prompt asking it to "probe like a senior researcher"; it is a testable interview contract. For each research topic, the contract should specify what researchers want to know, which concrete experience participants should address first, what can be neutrally probed if information is missing, what responses suffice, when to transition or stop, and which sensitive content must not be solicited. This preserves semi-structured flexibility while ensuring evidence across different conversations remains comparable.

AI can quickly generate the next question based on responses, but it is also more prone to embedding research assumptions into questions, repeatedly soliciting details, misunderstanding participants, or confidently filling in causal links never mentioned. The goal of the guide is not to make AI "chat like a human," but to restrict it to eliciting verifiable experiences grounded in what participants have offered, while preserving uncertainty, refusals, and counterexamples.

Separate Research Questions from Interview Questions

Research questions are unknowns the team needs to address, such as "what are the main decision pathways for first-time applicants abandoning the process"; they should not become direct questions like "why did you give up on your application?" The latter demands participants provide a complete causal summary in one sitting, which tends to elicit post-hoc rationalization. A better primary question is: "Starting with your most recent application attempt, what were you trying to accomplish? What happened next?"

The UK Government Service Design Manual also distinguishes between research questions for the team to answer and questions genuinely asked of participants, recommending that in-depth interviews use open, neutral starting questions focused on stories and real examples. Teams should first consolidate and prioritize questions from all stakeholders, then select those that interviews can answer; sales forecasts, incidence rates, and causal effects should not be silently included in scope merely because AI is used for interviews.

Structure Each Topic as a Six-Part Guide Contract

  1. Topic objective: What type of mechanism or evidence researchers hope to obtain; not revealed to participants.
  2. Entry condition: Which experience or response should trigger this topic; how to skip it when inapplicable.
  3. Primary question: Open, neutral language inviting a recent, concrete event.
  4. Evidence gaps: Missing elements such as time, people, actions, choices, reasons, outcomes, or verbatim language that warrant probing.
  5. Counterexamples and contrasts: How past successful experiences, alternatives, or people who did not face the issue differ.
  6. Completion and stopping: What information suffices before transitioning; when to end due to refusal, distress, sensitivity, or time constraints.

For instance, for a topic on abandonment, the objective is to reconstruct the path from the last task to the alternative behavior. The primary question asks, "When you last opened the service, what were you trying to do?" Evidence probes sequentially address what they saw, what they did, how they sought help, when they decided to stop, and what they used instead. If a participant says "I don't remember" or declines to share company details, accept the answer and move on; do not rephrase three times to keep pressing.

Use an "Evidence Packet" to Decide Whether to Keep Probing

A complete account of an experience should include at least context, actions, reasons, and outcomes. Context covers timing, goals, and constraints; actions describe what was actually done; reasons should stay as close to the participant's own understanding as possible; outcomes include completion, alternatives, and impacts. In each round, AI should only check for gaps in the evidence packet and ask one minimal probe, rather than running through all pre-written questions.

If a participant says "It was too much hassle," you can ask, "You mentioned hassle — at which step did it start?" If they say, "The upload kept erroring, so after two tries I switched to email," context, action, and alternative are already clear; the next question could focus on how the errors influenced the decision. You should not ask, "Was the upload button too hard to find?" because the button's placement has not been raised by the participant.

Classify Probes by Five Functions

  • Clarification: "Who do you mean by 'they'?" Resolves references and meanings only.
  • Process: "What did you do first after seeing the prompt?" Reconstructs sequence rather than asking for an overall evaluation.
  • Evidence: "Can you recall the gist of the prompt at the time?" Elicits verbatim expressions, allowing for imperfect recall.
  • Contrast: "What was different when it went smoothly before?" Helps distinguish stable obstacles from one-off events.
  • Confirmation: "So I understand you first tried re-uploading, and after it failed again, you switched to email, right?" Summarizes only what was said and allows correction.

The U.S. Census Bureau interviewer manual emphasizes neutral probes and stopping unnecessary probing once the response meets the question's objective. An AI guide likewise needs to specify "what not to ask": do not presuppose emotions, do not imply that one answer is more reasonable, do not create conformity with phrases like "many people feel," do not promote participant speculation to fact, and do not automatically fire multiple follow-ups after silence.

Mark Active Solicitation vs. Spontaneous Mentions in the Guide

If the team suspects price causes abandonment, you may include a neutral topic on decision factors across all interviews, but you cannot start with "Was the price too high?" A participant spontaneously mentioning price in their main narrative and choosing price after the researcher shows a price list are not equivalent evidence. The analysis table should record whether each topic arose spontaneously, via neutral probing, or after explicit prompting.

This distinction prevents miscalculating "mention rates." Eight of twelve people agreeing after being asked "Was the price too high?" cannot be written up as "two-thirds spontaneously abandoned because of price." Qualitative interviews are not suited for estimating population incidence anyway; prompted frequencies reflect guide coverage more than market proportions.

Pre-test with Twenty-Four Interviews: Comparing Two Versions of an AI Guide

Below is an example pre-test. The team recruited twenty-four participants who had recently discontinued the same application process and met identical screening criteria, randomly assigning them to an older question list or the new six-part contract, twelve each. Both versions had the same target duration, introduction, and five required topics: initial task, abandonment triggers, recovery attempts, alternatives, and final outcomes.

Five topics multiplied by twelve interviews give sixty "participant-topic" coverage opportunities per group. After independent review of raw transcripts by two researchers blinded to version using pre-defined rules, the old version covered 43 topics, or 71.7 percent; the new version covered 54, or 90 percent. Only four interviews in the old version obtained a complete evidence packet covering context, actions, reasons, and outcomes, versus nine in the new version. The old version included 18 instances of leading probes that embedded team assumptions in questions; the new version had 5.

These numbers support moving the new version to a small-scale field trial, but they do not prove it works better across all topics and populations. Since topic coverage units come from the same interviews, they are not independent; twenty-four people is also small. More important is to examine where the gaps cluster: if the new version fatigues participants, increases sensitive disclosures, or shortens answers in final topics in order to cover all five, then you need to reduce the number of topics or improve transitions, rather than simply chasing ninety percent.

Pre-testing Must Include Difficult Responses, Not Just Ideal Ones

The test set should include at least: one-word answers, long tangents, vague concepts, contradictions, poor recall, outright refusals, third-party privacy concerns, business secrets, unexpected sensitive disclosures, and disconnections. Check item by item whether AI asks only one question, whether it references content the participant never mentioned, whether it repeats itself, whether it respects refusals, and whether it completes key topics when time runs out.

NIST's generative AI risk management resource lists "confidently presenting incorrect content" as a confusion risk and recommends empirically validating capabilities, testing in environments close to deployment, and documenting human oversight and ongoing monitoring. For AI interviews, simulating a few smooth responses is far from sufficient; you must test the full conversation in the target language, on real devices, with varied expressions, and under high-stakes boundaries.

Time Budgets Should Be Aligned with Evidence Priorities

The guide may categorize topics as "must-cover," "conditional," and "optional." Start by confirming eligibility, consent, and the most recent event; in the middle phase, prioritize completing the most important evidence packets; only then proceed to comparisons and suggestions. Before each transition, AI should check: whether the current topic has enough information, whether the participant still has unfinished narratives, and whether remaining time supports a new topic.

Do not drag interviews out indefinitely simply because AI can keep asking. Longer interviews increase participant burden and may degrade data quality in the latter half. When a participant explicitly says "I don't know," do not treat it as a void to be conquered; "don't know" may itself indicate invisible information, an inapplicable role, or memory limits.

Sensitive Topics Require Specific Stopping and Escalation Rules

The guide should explicitly prohibit soliciting passwords, identification documents, account details, third-party identities, or unrelated health, financial, or work secrets. If participants show distress, crisis, potential harm, or complaints requiring formal handling, AI may only follow professionally reviewed scripts to stop, offer approved channels, or involve authorized personnel; it cannot improvise medical, legal, or psychological advice.

Even "I understand how you feel" may be overly humanizing or imply understanding. A safer expression is to acknowledge the participant has chosen to skip, and state that they may continue, rest, or end. Informed consent, recording, retention, and withdrawal rules must be governed separately from guide versions; a single prompt saying "be careful about privacy" cannot replace them.

From Raw Transcripts to Conclusions: AI Summaries Only as Indexes

Every conclusion should return to participants' exact words and their context. Automated summaries can miss negations, splice together experiences from different people, present speculation as fact, or homogenize wording. Researchers must personally review all high-risk segments and key topics, and retain guide versions, model or configuration versions, anomaly logs, and human correction records.

When publishing qualitative findings, at minimum describe the study population, recruitment and exclusion criteria, interview modality, dates, number of sessions, duration, guide version, analysis methods, and limitations. AAPOR's transparency disclosure elements also include the discussion script, semi-structured interview guide, and instructions to researchers, moderators, and participants as materials to be provided. Only then can readers judge whether conclusions arose spontaneously from participants or were explicitly prompted by the guide.

Pre-publication Checklist

  1. Research questions, decisions, and the boundaries of evidence interviews can answer are written clearly.
  2. Each topic includes objectives, entry conditions, primary questions, evidence gaps, counterexamples, completion, and stopping rules.
  3. Primary questions anchor on the most recent, concrete experience; they do not require participants to explain all causes at once.
  4. Probes address actual gaps in context, actions, reasons, or outcomes only.
  5. Spontaneous, neutrally probed, and explicitly prompted mentions are separately marked; prompted frequencies are not treated as incidence rates.
  6. One-word responses, tangents, contradictions, refusals, sensitive disclosures, and disconnections are pre-tested.
  7. Coverage, evidence completeness, leading, repetition, respect for refusals, and duration all have review metrics.
  8. Forbidden sensitive topics, stopping, escalation, and support scripts have undergone professional review.
  9. Guide versions and AI configurations in formal fieldwork are traceable, with major changes reported in batches.
  10. Automated summaries are verified against raw transcripts; counterexamples and uncertainties feed into conclusions.

Before choosing a method, you can read Exploratory, Descriptive, and Causal Research; when conversational media are involved, see Choosing Voice, Video, and Text Interviews; for sensitive studies, use Informed Consent and Privacy in AI Interviews; and after data collection, check open-text coding with Open-ended Survey Response Analysis. The benchmark for a good guide is not how many questions AI asks, but whether every conclusion can point to participant-provided evidence, a clear elicitation path, and respected boundaries.

Research Basis for Interview Design