GuíaMixed-Methods ResearchAI InterviewsQuantitative Surveys

How Do Questionnaires and AI Interviews Form a Genuine Mixed-Methods Study?

Using a 1,200-person questionnaire, 24 stratified purposive interviews, and a 600-person randomized validation, this article explains study sequencing, sample linkage, joint displays, conflicting evidence, and auditable inference.

Actualizado recientemente 2 de septiembre de 2026 27 min de lectura

Bottom line: Putting a questionnaire and AI interviews in the same project does not equal conducting mixed-methods research. A genuine mixed-methods study must align with a single decision, pre-specify what the quantitative and qualitative components each answer, why the sequence is ordered as it is, how the first phase reshapes the questions or sampling in the second, and where the two types of results converge to form new judgments. Questionnaires are better suited for estimating ranges, group differences, and uncertainty; interviews are better suited for reconstructing processes, language, mechanisms, and counterexamples. Interview themes cannot serve as population proportions, nor can questionnaire correlations be elevated to causality simply by appending a few quotes.

Three Basic Sequences Correspond to Three Types of Unknowns

An exploratory sequence is “interviews first, questionnaire later.” When the team is uncertain about key constructs, user language, or closed-ended options, they first build candidate mechanisms from diverse cases, translate them into neutral question items with cognitive pretesting, and then use the questionnaire to estimate distributions. The first phase generates hypotheses and measurement content, not population prevalence rates.

An explanatory sequence is “questionnaire first, interviews later.” When a stable difference has emerged from the questionnaire but the mechanism is unknown, interviews are purposively invited based on subgroups, extremes, typical cases, and counterexamples. The second phase explains quantitative patterns and may also reveal item misinterpretations or unobserved confounders. A convergent parallel design has both types of data answering shared questions within a similar timeframe, followed by comparisons of confirmation, complementarity, and dissonance; it is faster, but if there is no pre-specified shared construct and joint display framework, it can easily produce two unrelated reports.

Interviews can also be embedded within randomized trials to explain why an intervention works for some individuals and not others. Complexity is not inherently better. If the decision only requires an overall satisfaction rate, adding interviews may be an unnecessary cost; if the task is simply to uncover process barriers, a large representative questionnaire may not be the right first step. The justification for a mixed-methods approach should be that a single method indeed leaves a critical evidence gap.

A 1,200-Person Questionnaire Detected a Difference, but Could Not Explain the Mechanism

A enterprise software team aimed to improve first-time configuration completion rates among administrators. The survey linked 1,200 administrators with genuine configuration tasks during the same observation window to their task status: among 400 new administrators, 156 did not complete, a failure rate of 39%; among 800 experienced administrators, 168 did not complete, a failure rate of 21%; overall failure was 324 ÷ 1200 = 27%, with new administrators 18 percentage points higher than experienced counterparts. This difference can prioritize target groups but does not explain whether experience alters knowledge, confidence, organizational authorization, or task complexity.

In the same survey, 83% of new administrators agreed with the statement “I understand the meaning of each permission,” yet actual completion rates were only 61%. The discrepancy between attitude and behavior may stem from acquiescence bias, overly broad items, barriers between understanding and execution, or flawed log data. The report should not select the more favorable 83% as the success metric, nor can it conclude causally from correlation alone that “lack of permission knowledge causes failure.” The next phase needed to specifically explain this contradiction.

The 24 AI Interviews Must Be Rule-Based from the First Phase

At the end of the questionnaire, the team separately obtained consent for follow-up contact, separating contact details from survey response data and linking them via controlled tokens. Interviews were not selected based on convenience but followed a four-cell sampling frame: 8 new administrators who failed, 4 new administrators who succeeded, 8 experienced administrators who failed, and 4 experienced administrators who succeeded, totaling 24. Failure cases were intentionally oversampled; experience and outcomes were both controlled. Selection rules, including invitation, refusal, and completion counts per cell, were recorded in the sample flow.

Among the 16 failure cases, 12 could not anticipate which data a specific role would see; among the 8 success cases, 6 relied on legacy tables, colleague templates, or trial-and-error across accounts. These numbers only describe these 24 purposively sampled interviews and cannot be generalized as “75% of failed administrators do not understand permissions” or “75% of successful ones rely on templates.” The genuinely new information they provide is that the abstract “understanding permissions” self-report conceals the ability to “preview actual access outcomes,” and that the advantage of experience may partially stem from workarounds developed outside the organization.

Researchers should also actively search for counterexamples. Four failed administrators could accurately explain permissions but stopped due to lack of approval rights; two successful ones had no templates but faced simpler tasks. These cases break the original hypothesis into at least three mechanisms: visibility of permission consequences, reusability of configuration, and organizational approval. If only low-scoring complainers are recruited, the team may attribute all failures to interface issues.

Use Joint Displays for Analysis, Not Merely Placing Two Chapters Side by Side

Before data collection, a five-column joint display should be built: quantitative pattern, qualitative evidence, relationship, integrated inference, and next-step verification. The first row might read “new administrators failed 39%, experienced 21%”; the interview column notes “successful ones often used templates; failures found it hard to preview data access”; the relationship is labeled “extension, not proof”; the integrated inference is that “experience differences are consistent with preview and reuse support, but interviews cannot establish causality”; the action is to test role preview and templates.

A second row places “83% self-reported understanding, but 61% actual completion”; interviews reveal some could repeat labels yet cannot judge specific account outcomes; the two types of evidence diverge. The integrated inference is not to average, but to recognize that the existing self-report item may lack validity, and future measurement should shift to situational judgment or real tasks. A third row records the approval-rights counterexamples, reminding that a product redesign will not solve all failures. Joint displays force every conclusion to simultaneously address scope, mechanism, and boundaries.

Carry the Integrated Inference into a 600-Person Randomized Validation

Based on this, the team developed a “role preview plus reusable template” and randomized 600 new administrators: in the current process, 186 of 300 completed, a 62% completion rate; in the new process, 222 of 300 completed, a 74% completion rate. The absolute difference is 12 percentage points, with a common independent-proportion approximate 95% confidence interval of about 4.6 to 19.4 percentage points. The randomized validation makes the causal claim “the new process increases completion rates” more credible than the first-phase cross-sectional difference, but it still only covers this population, task, and implementation version.

Interview mechanisms cannot be declared fully confirmed simply because the trial succeeded: the new process included both preview and templates, making it impossible to isolate their individual contributions. Next steps could include factorial designs or staggered rollouts, along with monitoring guardrails such as misauthorizations, configuration reversals, completion time, and support requests. The role of mixed methods is to convert “who fails” into testable mechanisms and actionable changes; final effectiveness must still be answered by an appropriate validation design.

Linking Samples Must Protect Both Inference and Privacy

Two phases should share core eligibility, role definitions, and time windows; otherwise, interviews may be explaining a different population. When linking individuals, retain only the minimal identifiers needed for sampling. Store contact details separately, restrict access, and set deletion times. The invitation should specify the interview medium, whether AI moderates or transcribes, the purpose of recordings, duration, compensation, and how to withdraw. Declining the second phase should not affect the legitimate processing of first-phase data; specific rules follow consent statements and applicable requirements.

Those who agree to follow-up interviews often differ from those who decline. Report the number of questionnaire respondents who consented to contact, the number drawn, actual invitations, and completions, and compare key known characteristics; this does not eliminate self-selection bias but informs readers about whom the interviews represent. For sensitive studies, assess whether linking questionnaire responses to interview quotes increases re-identification risk, and remove unnecessary combinations of features from reported quotes.

Additional Methodological Disclosures Are Required for AI Interviews

Method appendices should archive questionnaires, interview guides, AI moderator instructions, probing boundaries, model or system versions, voice or text modes, language, quality checks, and human review procedures. Auto-generated summaries must be verified back to verbatim transcripts or original text; model-generated transition sentences cannot be presented as participant quotes. If probing depth varies systematically by accent, utterance length, or sensitive topics, spot-check by subgroup rather than reporting only overall completion counts.

Questionnaires and interviews should each be analyzed according to their respective standards: quantitative sections report population, sample, weights, bases, missingness, and intervals; qualitative sections report purposive sampling, analytical units, coding processes, counterexamples, and researcher judgments. Finally, indicate who performed the integration and where it occurred in sampling, instrument development, analysis, or interpretation. Tool integration can reduce field loss but cannot replace methodological integration.

Conflicting Evidence Is a Product, Not a Failure

Two types of results may corroborate, extend, contradict, or remain silent. Corroboration indicates that different evidence supports converging judgments; extension means one type provides scope, and the other adds mechanisms; contradiction may arise from samples, timing, item interpretation, social desirability, mode effects, or genuine heterogeneity; silence indicates that a method did not cover an issue identified by the other method. Pre-specify conflict resolution processes: re-examine definitions and source material, check sample differences, solicit counterexamples, conduct minimal validation if necessary, rather than allowing team leads to choose whichever side aligns with prior expectations.

Research planning can begin by referencing Exploratory, Descriptive, and Causal Research to choose the sequence; maintain consistent definitions between phases using Defining Target Survey Audiences; control probing in the interview phase according to AI Interview Guide Design; and present integrated inferences via Actionable Insight Reports. The most valuable outcome of mixed methods is not more data but placing scope, mechanisms, and verifiable actions onto the same chain of evidence.

References on Mixed Methods and AI Interviews