GuideAI interviewsvideo interviewsmode effectsvoice interviews

AI Interview: Text, Voice, or Video? Decide by Evidence Yield, Not Richness

Compare the three modalities on evidence capability, completion burden, mode effects, transcription errors, privacy, and accessibility, then select a single or mixed mode based on the usable evidence per invitee.

Mis à jour récemment 2 septembre 2026 30 min de lecture

AI interview modality should be chosen as the least burdensome option that still obtains the evidence required for the study — not defaulting to video as the richest, voice as the most natural, or text as the safest. Text suits asynchronous review and precise wording; voice suits continuous narration; video is necessary only for tasks where actions, objects, and spaces must be observed. But the three modes change who can participate, how answers are given, social pressure, technical failure rates, and privacy exposure. Because the modes differ, results cannot be merged directly into a single frequency count.

Before deciding, write one sentence: "Without seeing ____, we cannot answer the research question." Only if the answer is hand movements, package opening, screen state, or spatial layout is there sufficient justification to request video. If the final decision process is all that is needed, voice or text is usually sufficient. Collecting extra faces, backgrounds, and sounds does not automatically improve insight—it expands data governance and dropout costs.

What Evidence Each Modality Is Best At

  • Text: Participants can think, review, and revise asynchronously; raw content is easy to search and quote; suitable for contexts where speaking is inconvenient. Downsides: high input burden, spelling/reading/keyboard barriers shorten responses, and rhythm and tone cues are limited.
  • Voice: Good for narrating timelines, turning points, direct quotes, and emotional shifts; hands are free from typing. Downsides: requires a relatively private and quiet setting; accent, noise, overlapping speech, and proper nouns increase transcription errors.
  • Video: Can observe actions, objects, environment, and screen; allows comparing "what is said" versus "what is done." Downsides: highest bandwidth, permission, appearance pressure, and third-party exposure; analysis must separate observation, interpretation, and researcher inference.

"More evidence" is not the same as "better evidence." When studying package opening, the action in video may be core data. When studying work difficulties, asking for an office view may expose colleagues, documents, and business information without adding needed evidence. For each piece of visual information, specify its purpose; if there is no purpose, ask participants to turn off the camera or decline sharing.

Six Questions to Decide the Mode

  1. Must the conclusion rely on observing actions, objects, interface, or environment?
  2. Do participants need to narrate continuously, or do they need time to look up records and carefully word their answers?
  3. Is the topic sensitive, does a human presence, voice, or face increase social pressure?
  4. What are the target population's Internet access, devices, literacy, hearing, speaking, vision, and fine-motor conditions?
  5. How is the accuracy of transcription, translation, captioning, and descriptive notes verified?
  6. Does the team have the necessary permissions, secure spaces, and retention rules for handling raw audio/video?

If any single mode would systematically exclude a segment of the target population, provide a real alternative path, or explicitly restrict the study scope to those who can use that mode. Letting participants "choose freely" helps coverage, but the reason for choice is confounded with outcomes; if the goal is to compare modes themselves, random assignment is needed within ethical and accessibility limits, along with records of refusal and switching.

A 360-Person Case: Video Was Richest, but Not the Most Efficient

Below is an illustrative pretest. The team studied a non-sensitive online service interruption experience. From the same sampling frame, they drew three hundred and sixty eligible people and randomly invited them to text, voice, or video groups (n=120 each). Topic, incentive, time window, and outline objectives were identical. Participants could withdraw, but to measure modality burden, no switching was allowed.

Text group: 108 completed (90%); voice: 96 completed (80%); video: 78 completed (65%). Researchers pre-defined an "evidence package" as containing specific events, participant actions, reasons, and outcomes; two coders unaware of mode goals rechecked. Qualifying packages: 72 text, 78 voice, 66 video.

Among those who completed the interview, package rates were 66.7%, 81.3%, and 84.6%, with video highest. However, with every 120 invitees as the denominator, usable evidence yields were 60%, 65%, and 55%, respectively. Video was richer among completers but, because more people did not finish, had the lowest evidence yield per invitee; in this scenario, voice was the better primary mode.

This result is not a universal ranking. The topic needed no visual evidence, so video added little information; if the task were observing package use, the text evidence definition would be different. Completers may also be those who are more comfortable with video or have better equipment. The correct conclusion is: "Under this population, topic, invitation, and technical conditions, voice obtained the highest usable evidence per invitation." The next round should still offer a text path for those who cannot or will not use voice.

Mode Changes Answers, Not Just Length

Social interaction in voice or video may make participants more polite; text self-administration may allow more direct expressions on sensitive topics but also produces shorter, more edited narratives. Pew Research Center once compared telephone interviewer to web self-administration in a large randomized experiment: many questions showed no mode difference, but some did, and larger differences often involved social desirability. That study is not an AI in-depth interview experiment, but it clearly shows mode differences cannot all be treated as attitude differences across populations.

When analyzing mixed modes, keep mode, reason for choice, completion, duration, word count, technical failures, and evidence package quality. Use identical definitions for core themes; report mode-specific evidence separately. Do not infer deeper engagement just because voice transcripts are longer, and do not interpret silence in text as lack of experience.

Speech Transcription Must Be Evaluated in Real Listening Conditions

Transcription tests should not be limited to team members reading standard sentences in a quiet meeting room. Tests should cover target languages, accents, speaking rates, pauses, code-switching, proper nouns, background noise, phone microphones, and talking while doing tasks. NIST's early speech recognition evaluation materials already point out that read-speech recognition can outperform concurrent-task and natural settings; reports need to describe the speaking task and environment.

Quality checks can sample stratified segments to separately inspect key nouns, negation, numbers, speakers, temporal order, and inaudible markers. When the machine is unsure, it should mark uncertainty, not let the language model fill gaps based on context. Quotes, risk events, and decisive segments should be rechecked against the original audio; if consent does not allow retaining raw audio, perform necessary verification before deletion and document limitations.

Video Analysis: Separate What You See, Hear, and Infer

Video can document "the participant rotated the package three times before finding the opening," but "they were frustrated" is an interpretation unless the participant says so or a predefined behavioral rule applies. Facial expressions cannot automatically be treated as true emotions, personality, or intentions. Screen sharing may expose contacts, messages, addresses, client data, and passwords—invitations and interfaces should remind users to close notifications, use sample data, and stop sharing anytime.

Cameras also capture family members, colleagues, and home environments—third parties who have not necessarily consented. When research needs only hands or products, allow participants to adjust the camera to avoid faces and backgrounds; when research needs only the screen, do not simultaneously collect video. If unrelated sensitive information is discovered, follow a preset process to restrict access, edit, or delete—rather than keeping it "just in case."

Accessibility Is Not Forcing One Mode on Everyone

UK government remote research guidance notes that phone and video can help some people with mobility or reluctance to attend in-person, but may exclude Deaf people and those with speech impairments and create privacy and technical issues. Truly inclusive design offers text, captions, sign language, or support persons as appropriate for the target population, allows more time, and pretests with participants' own assistive technologies.

W3C Web Accessibility Initiative states that captions must provide speech and non-speech audio information needed to understand content; descriptive transcripts also include key visual information. Therefore, an auto-generated transcript with only dialogue does not make a video accessible, nor does it give analysts who did not watch the video equivalent evidence. For visual tasks, a structured observation record should accompany the video.

Two Designs for Mixed Modes: Do Not Collapse Them

Coverage-based mixing lets participants choose the most suitable mode, aiming to reduce exclusion; analysis must acknowledge self-selection differences. Comparison-based mixing randomly assigns mode where feasible, aiming to estimate how mode changes completion or answers; it requires identical invitation, content, time, and outcome definitions. The former cannot directly estimate mode effects; the latter should not force someone with accessibility needs into an inappropriate mode.

Alternatives can use a primary mode plus escalation path: start with text or voice retrospective, and only when a key research question truly requires observation, ask for a separate short video task with its own consent. This limits extra data collection to participants and segments where it has clear value, rather than making everyone default to video.

Privacy, Consent, and Retention Should Be Specified by Modality

Invitations must state whether text, voice, face, screen, or environment is collected; which part of the AI pipeline uses what (e.g., follow-up, transcription, analysis); who has access; how long it is retained; and whether participants can turn off the camera, stop sharing, skip, or exit. Audio, video, and transcripts are separate data objects; agreeing to an "interview" does not imply consent for all raw data to be kept long-term.

If analysis needs only text, consider deleting raw media after accuracy checks; if motion evidence must be kept, limit segments, access, and duration. Specific plans must comply with local law, institutional ethics, contracts, and security requirements; this article does not replace professional review.

Checklist Before Publishing

  1. Every audio or visual evidence item has an explicit research purpose.
  2. The chosen mode is the least burdensome, and a real alternative exists for those who cannot use it.
  3. Completion rate, quality among completers, and usable evidence per invitee are calculated separately.
  4. Self-selected modes are used for coverage; random modes are used for comparison; these conclusions are not blended.
  5. Core themes are consistent; mode-specific evidence and raw word counts are reported separately.
  6. Transcription was pretested under real accents, noise, devices, and concurrent tasks.
  7. Key negations, numbers, quotes, and risk segments are verified against the original media.
  8. Video observations, participant interpretations, and researcher inferences are labeled separately.
  9. Captions, descriptive transcripts, assistive technology, and participation time meet the target population's needs.
  10. Text, audio, video, screen, and transcripts are each subject to separate information, authorization, access, and deletion rules.

Dialogue content can be designed with AI interview guide methods; specific data boundaries should consider informed consent and privacy; pretesting can reference cognitive pretesting; thematic analysis of mixed-mode data should use open-text coding while retaining the mode field. The ultimate standard for choosing a mode is not how rich the session looks, but how many people in the target population can safely complete it, whether evidence answers the question, and whether key content can be accurately rechecked.

Sources on Modality & Accessibility