GuiaAI interviewInformed consentPrivacy protectionData minimization

How to Conduct Informed Consent and Protect Privacy in AI Interviews? Start with Data Flows, Not Checkboxes

Implement layered disclosure covering AI roles, recording media, purposes, access, retention, sharing, withdrawal, and incidental disclosure. Support commitments through object-level data inventories, minimization, and re-identification review.

Atualizado recentemente 2 de setembro de 2026 37 min de leitura

Informed consent for AI interviews is not a one-time checkbox before commencement—it empowers participants to anticipate, across the entire data lifecycle: who is conducting the research, why they are invited, what AI does, what will be recorded, who can access it, where data will go, how long it will be retained, how to skip or withdraw, and which individual-level signals will be processed. Disclosure without data minimization, access controls, deletion, and incidental disclosure procedures leaves protection at the level of rhetoric; security measures without comprehension and voluntary choice do not constitute adequate consent either.

Legal, ethical, and contractual requirements vary by jurisdiction, institution, and research type. Projects involving employees, minors, healthcare, finance, or cross-border transfers may incur additional rules. This article provides a research design framework, not specific legal conclusions. Before launch, authorized legal, ethical, privacy, security, and domain leads should review against current data flows and local requirements.

Map the Data Flow Before Drafting Consent Language

Teams should enumerate each data object from invitation to deletion, rather than stating generically, “We will protect your information.” A single AI voice interview may concurrently generate: recruitment lists, unique invitation links, eligibility responses, consent records, raw audio, automatic transcripts, conversation context, model inputs/outputs, human annotations, thematic coding, quotation excerpts, incentive records, follow-up contact details, aggregate reports, and system logs.

For each object, answer at least eight questions: Why is it needed, which fields are collected, from whom is it sourced, through which systems or vendors does it pass, who can access it, how is it linked, how long is it retained, and when is it deleted. If the team cannot articulate whether raw audio is retained after transcription, they cannot make truthful commitments to participants; if vendor configurations allow content to be used for model training or product improvement, a blanket statement like “AI is only used to assist analysis” is insufficient.

Short-Form Disclosure Should Answer at Least Ten Key Questions

  1. Who is the research entity and the contactable principal investigator?
  2. What is the research purpose, why is the participant invited, and what will they be asked to do?
  3. What is the expected duration, and what is the scope of text, voice, video, screen, or other recordings?
  4. Is AI responsible for asking questions, probing, transcribing, translating, summarizing, coding, or quality checks?
  5. What information is explicitly not needed, and what should participants avoid disclosing?
  6. Who accesses and receives raw recordings, transcripts, analyses, and quotations, respectively?
  7. Does the data pass through external services, cross-border processing, or use for training and improving models?
  8. How long is each data category retained, how is it deleted, and what are the practical limitations of deletion?
  9. Is participation voluntary, and can participants skip, pause, or withdraw without consequences?
  10. What channels are available for questions, privacy requests, complaints, risk, or emergencies?

The U.S. Department of Health and Human Services Office for Human Research Protections describes consent as a process of disclosing information, fostering comprehension, and maintaining voluntariness, clarifying that a signed document itself does not equate to adequate consent. A practical approach is to present a one-screen summary of key points first, followed by full details; ask someone unfamiliar with the project to articulate the AI’s roles, access, retention, and withdrawal procedures rather than merely checking whether they clicked “agree.”

Anonymity, De-identification, Pseudonymization, and Confidentiality Are Not Interchangeable

Anonymity means the research team cannot reasonably link data back to individuals; projects with unique invitations, incentives, or follow-ups rarely achieve true anonymity at the collection stage. De-identification involves removing or generalizing direct identifiers such as names, but re-identification may still be possible given context. Pseudonymization replaces direct identities with codes, but a linking table still exists. Confidentiality limits disclosure through commitments, permissions, and processes.

Therefore, do not describe a survey as unlinkable to individuals. A more accurate formulation might be: “Invitation lists are stored separately from interview content; analysis copies remove direct identifiers; only designated researchers use the linking key when necessary; administrators receive only aggregated results meeting reporting thresholds.” Protection capabilities must align with actual system configurations.

The 120-Employee Case: Re-identification Without Asking for Names

The following is an illustrative case. An organization invited 120 employees to discuss cross-team collaboration, with 20 from each of six departments. The study did not ask for names during the interviews, but the recruitment list retained department, job level, and tenure. Eighteen volunteer follow-up participants left contact information in a separate form. The initial draft report cross-tabulated results by department, job level, and “tenure less than one year.”

In one department, the cell for “senior role and tenure less than one year” contained only three individuals; one quotation referred to a migration project managed by only one member. Even without names, colleagues could easily infer the speaker from department, tenure, role, and project context. The risk stems from field combinations and verbatim quotes, not solely from a name field.

The team consequently eliminated independent proportions for that three-person cell, merged tenure into broader ranges, and removed project names that did not affect conclusions; quotations were converted to paraphrases of common mechanisms, using verbatim quotes only when patterns recurred across larger populations and could not be reasonably identified. The 18 follow-up contact records were retained in a more strictly access-controlled separate table; the report analysis table stored only random codes. Even with minimal reporting units, this is not an anonymity guarantee; unique events and organizational knowledge still require case-by-case review.

Data Minimization Must Be Implemented Across Questioning, Media, and Retention

Minimization in questioning: For each sensitive field, specify which analysis or decision it informs. When only company size is needed, do not ask for the full company name; when only region is needed, do not collect precise addresses. Passwords, verification codes, full identification numbers, full account details, third-party credentials, and unrelated health or financial information should be explicitly prohibited.

Minimization in media: Do not default to recording voice and face when text evidence suffices; do not capture home environments when only observing hand movements during a screen-share task. Before screen sharing, remind participants to close notifications and use sample data. When AI outlines encounter sensitive content, do not probe further; transition according to approved rules.

Minimization in retention: Raw media, transcripts, linking keys, coding, and aggregated results have different retention periods. The UK Information Commissioner’s Office guidance on AI and data protection identifies purpose, retention, sharing, use limitation, and minimization as key issues, noting that “might be useful later” does not automatically justify current collection or retention. This guidance is being updated in line with UK legal changes, and specific obligations must be verified against the latest applicable rules—but the methodological principle of preventing function creep remains clear.

Articulate Each AI Role Explicitly

“This study uses AI” is insufficient information. Participants need to know whether AI generates follow-up questions based on the outline, transcribes speech, translates, summarizes, recommends topics, or produces outputs that may affect individual assessments. They must also be informed whether researchers review raw content, whether automated outputs are human-reviewed, how errors are corrected, and whether conversations are used to train or improve external or internal models.

NIST’s Generative AI Profile materials recommend documenting and informing users about AI interactions before significant interactions; system records should include data provenance, known issues, human oversight, sensitive data, and underlying model versions. For interview projects, this means “consent version—outline version—model or service version—data usage” should be traceable, not merely a timestamp of checkbox acceptance.

A Changed Purpose Cannot Be Indefinitely Covered by Old Consent

If the initial purpose was “summarize feedback for this service improvement,” but later the raw conversations are intended for training general models, employee performance evaluation, sales leads, or product qualification judgments, this exceeds the participant’s reasonable expectations. The team must pause the new use, conduct fresh legal and ethical assessments, determine compatibility, and obtain renewed disclosure or consent as needed; if conditions are not met, the data must not be used.

Do not attempt to stretch authorization indefinitely with a clause at the end of a long document stating “may also be used for other purposes.” Disclosure must be specific enough for participants to understand why they might or might not be willing to proceed. When there are substantive changes in purpose, sharing partners, retention periods, or risks, document new versions and their applicable populations.

Withdrawal and Deletion Must Describe Realistic, Enforceable Boundaries

Participants should know how to pause, skip, or end the interview, and by what point after the interview they can request deletion. If content has already been irreversibly included in aggregate statistics that are disconnected from identifiers, or if certain records must be retained by law, state this clearly in advance. Do not promise “delete all your data at any time” when there is no actual mechanism to locate relevant copies.

The deletion process must cover primary storage, exports, transcripts, annotations, backup strategies, and vendor-side disposal, with documented completion. The research team should also decide whether withdrawal affects reasonable incentives already earned; do not use incentive forfeiture to punish mid-study withdrawal, as this undermines voluntariness.

Employees, Children, and High-Risk Contexts Require Additional Safeguards

Even when employees see “voluntary,” they may worry that their direct supervisor will know who declined or what they said. Recruitment and analysis should be separated from performance management; managers should not see individual responses; reports should control small cells and identifiable quotes. Lotteries, work time, or supervisor invitations should also be assessed for undue pressure.

When involving children, obtaining guardian consent alone is insufficient—children’s own age-appropriate information and assent must be addressed. UNICEF’s evidence ethics policy emphasizes ethical review, protection, risk management, meaningful consent or assent, confidentiality, privacy, and data management. Children’s refusal or cessation must be respected; specific practices are determined by local law, age, capacity, institution, and risk.

For contexts involving trauma, medical care, financial hardship, illegal activity, or safety incidents, designated professionals should establish protocols for stopping, support, and escalation. AI must not make promises about confidentiality exceptions, emergency responses, or professional advice on its own. Participants must know whether someone is monitoring the questionnaire or interview in real time and the official channels for emergencies.

Open Text and Automated Summaries Present Two Privacy Risks

The first risk arises at the point of raw disclosure: participants may voluntarily reveal names, companies, clients, addresses, accounts, or third-party events. Systems should prompt users to avoid unrelated information, implement safety blocks for high-risk patterns, and allow participants to review or correct responses. The second risk occurs in summarization: models may retain unique details or combine multiple threads into a more identifiable profile.

De-identification cannot be a one-time name replacement. Check for direct identifiers, quasi-identifiers, unique events, quotation necessity, and small cells; human-review all external-facing verbatim quotes. Automated summaries serve as analytical indexes only; smoother phrasing cannot override the original consent scope.

An Auditable Data Inventory

The project inventory can be structured as rows for each “data object” and columns for purpose, fields, sources, legal/ethical basis, systems, vendors, regions, access roles, linking keys, retention periods, deletion methods, consent versions, and risk owners. The pre-launch review is not “Do we have a privacy policy?” but rather a line-by-line confirmation that disclosure matches actual processing.

Regularly audit permissions, anomalous exports, missed deletions, model or vendor changes, and participant requests. When deviations occur, document the scope of impact, remediation, and whether notification is required. Reports should only output the minimum aggregated results necessary for decisions; do not provide all raw data to every project member by default.

Pre-Launch Checklist

  1. Data flows for invitations, identity, consent, media, transcription, models, coding, quotations, follow-up, and deletion have been mapped.
  2. Short-form key points allow unfamiliar readers to articulate AI roles, access, retention, and withdrawal procedures.
  3. Terms such as anonymity, de-identification, pseudonymization, and confidentiality align with actual capabilities.
  4. Each field, medium, and retention period has a specific research purpose; if none, delete it.
  5. Vendor arrangements, cross-border processing, model training, and product improvements are accurately disclosed.
  6. Contact information, linking keys, and content are separated; access follows least privilege.
  7. Small cells, quasi-identifiers, and unique quotations have undergone re-identification risk review.
  8. Refusals, withdrawals, deletions, complaints, sensitive disclosures, and emergency support are all operational.
  9. New purposes and significant configuration changes trigger reassessment; old consent does not cover indefinitely.
  10. Consent, outline, model, processing, and reporting versions are aligned and auditable.

Media data objects can be paired with guidelines on text, voice, and video interviews; sensitive topics and stopping conditions should be embedded in the AI interview outline; database-layer separation of keys, identities, and responses should reference survey database schema design; versioning and deletion logs can leverage data versioning and lineage. The most reliable consent language is not the most comprehensive—it is the language that can be substantiated at every point in the actual systems, permissions, and deletion practices.

Consent and Privacy Sources