LeitfadenOpen-ended analysisText codingAI-assisted analysis

How to Code, Validate, and Turn Open-Ended Responses into Reliable Conclusions?

From analytical paradigms, sampling for codebook development, and dual coding to multi-label proportions, AI assistance, and rechecking original text, a complete worked example using 600 responses illustrates how open-ended responses can lead to auditable conclusions.

Zuletzt aktualisiert 2. September 2026 25 Min. Lesezeit

Bottom Line: Analyzing open-ended responses is not about throwing a few hundred responses at a model and asking for a summary. It starts with deciding which type of analysis you need, ensuring each theme can be traced back to clear code definitions, the appropriate denominator, and the original text. If your goal is to estimate “how many responses mentioned what,” use a tested codebook-based content analysis. If your goal is to explain how people make sense of an experience, consider thematic analysis which emphasizes researcher reflection, context, and contradictions. Both can yield insights, but their rules of evidence differ; you cannot claim to be exploring meaning while treating vague automatic labels as precise proportions.

Choose Your Analytical Paradigm Before Choosing Tools

Common tasks can be broken down into three types. First is retrieval: finding responses related to refunds, delivery delays, or safety incidents; the core is whether recall is sufficient. Second is codebook analysis: applying one or more labels to text according to stable rules to compare thematic prevalence across groups; the core is definition, reproducibility, and denominator. Third is interpretive thematic analysis: searching for meaning, relationships, and exceptions behind the statements; here, the researcher's judgment is part of the analysis process, with the core being reflection, argumentation, and context. Word frequencies, sentiment scores, and summaries are auxiliary outputs and cannot automatically replace any of these approaches.

For example, a response like “Customer service was polite, but I had to repeat my issue across three different channels, and it took seven days to get a refund” contains at least three codes simultaneously: positive attitude, cross-channel information fragmentation, and delay. A single sentiment label would compress a clear process failure into “neutral,” and merely counting the word “customer service” wouldn't reveal who should take action. Before starting, formulate your research question as a testable statement, such as “Which fixable factors are causing refund experiences to deteriorate?” instead of vaguely asking “What did users say?”

Create an Auditable Text Handling Ledger

Store the original text in a read-only format. Maintain separate fields for cleaned text, exclusion flags, reasons for exclusion, language, sample strata, codebook version, and coder ID. Blank entries, copied question text, unreadable characters, and completely off-topic responses can be excluded from thematic denominators, but the count must be reported. Highly similar responses shouldn't be deleted just because they “look like spam”; instead, assess them alongside response time, device repetition, and other quality signals. Obscure unnecessary identifiers like names, phone numbers, and order IDs before sharing with analysts or external models. Preserve the original records and processing maps in a controlled environment.

When drawing a random sample for codebook development, taking just the first few responses can bias your codebook with timing and channel order. A more robust approach is to ensure strata coverage for major population segments, languages, varied answer lengths, positive and negative experiences, and a few potentially anomalous items. If small but critical groups are decision-relevant, you may intentionally oversample them, but this means the code-development sample cannot be used to estimate overall proportions. The analysis ledger should detail: total record count, number of valid open-ended responses, how the code-development sample was drawn, which texts had dual coding, and which records were affected by each rule change.

The Codebook Should Answer Six Questions at Minimum

Every code should have a name, a business definition, inclusion criteria, exclusion criteria, examples, and boundary cases, plus specify the coding unit. The unit can be the entire response or a segment expressing a complete idea. Units too small can fragment negations and causality; units too large can force together multiple opposing sentiments. Multi-label coding is usually appropriate for open-ended questions: a single response can relate to process, price, and communication simultaneously. Don't force coders to choose just one “main” theme unless the research question genuinely requires mutually exclusive categories.

Clearly define the hierarchy and relationships between codes. For example, “experienced excessive delay” is a higher-order issue, while “initial response was slow” and “refund arrival was slow” represent different stages. Avoid using “bad attitude” as a catch-all for all dissatisfaction. Maintain an “other – needs review” category and an analytical memo to capture new phenomena not covered by the initial codebook. Use version numbers when codes change: if you split “slow response” into two sub-codes, reprocess all affected old records rather than having early and late batches use different definitions.

A Complete Worked Example with 600 Customer Service Responses

Suppose a post-purchase survey yields 600 valid open-ended responses. The team first stratifies by customer type, channel, and text length, selecting 120 responses to build the codebook. Then, from responses not used in that development, 80 are drawn and two coders independently judge whether “support response delay” is present. Coder A marks it present in 34 responses; Coder B marks it present in 30. Both agree it is present in 26, both agree it is absent in 42, and they disagree on 12.

The simple percentage agreement is (26+42)/80 = 85.0%. However, even if the two coders randomly judged according to their own overall positive rates, some agreement would occur by chance. Coder A's positive rate is 34/80 = 42.5%, Coder B's is 30/80 = 37.5%. The expected agreement by chance is (42.5% × 37.5%) + (57.5% × 62.5%) = 51.9%. Thus, Cohen's κ in this example is (85.0% - 51.9%) / (1 - 51.9%) ≈ 0.69. This figure is not a “certificate of quality”: κ is influenced by how common the code is (κ is lower for rare codes), and it can change with different coding units. The team must then scrutinize the 12 disagreements. They discover it was unclear whether “I got an automated reply but no one resolved it for two days” counted as an initial response delay. So, they refine the code to “time to get a response from a human who could resolve the issue exceeded the promised timeframe” and retest on another sample.

After freezing the codebook and coding all 600 responses, they find 198 responses mention response delay, representing 33.0% of valid responses. 144 responses mention an unclear answer, representing 24.0%. Of these, 96 responses mention both, representing 16.0%. Because this is multi-label coding, 33.0% and 24.0% cannot be summed to a mutually exclusive 57.0%. More action-oriented conclusions are: among the 198 responses mentioning delay, 96/198 = 48.5% also mention unclear information. This suggests that simply shortening initial response time may be insufficient; reply quality needs to be part of the improvement plan.

What Proportions Can and Cannot Tell You

The denominator for open-ended theme proportions could be all survey completers, all those who answered the open-ended question, or only those who experienced a certain situation. For instance, if the full survey had 1,000 completers but only 600 answered the open-ended question, “response delay” could be reported as 33.0% of valid open-ended responses or as 19.8% of all completers. These two numbers answer different questions. Not mentioning something doesn't mean it didn't happen. Question wording, open-ended position, mobile input burden, and respondent motivation all influence appearance rates, so don't directly equate theme proportion with issue prevalence.

When comparing groups, also consider unweighted base sizes and sampling design. If there are only 25 valid texts from enterprise customers and 12 mention delay, that surface-level 48% shouldn't be ranked with the same precision as several hundred consumer responses. For extremely rare but severe issues like safety, discrimination, or compliance violations, establish an independent escalation path so they aren't lost outside the “top five themes” due to low frequency. Quotations are meant to illustrate language and mechanisms, not to prove prevalence; before publishing, ensure they are de-identified and cannot be combined to reveal an individual.

AI Can Scale Processing, But Cannot Replace the Evidence Chain

AI is suitable for suggesting candidate codes, pre-labeling full datasets against a frozen codebook, clustering similar phrasing, aiding translation, and locating counterexamples. People should be responsible for determining the unit of analysis and paradigm, checking code meaning, validating key proportions, and reading original text. Generative models may invent causal links, tone, or quotations not present in the source text, so every representative snippet you report must link back to an actual record. Model-generated “typical user quotes” cannot be presented as participant evidence.

Don't just validate on a simple random sample, because high-frequency codes can mask failures in rare categories. Combine random sampling with sampling per key code, low-confidence samples, and oversampling of sensitive groups, reporting error separately for each. If the model version, prompt, or codebook changes, freeze the configuration and rerun the affected parts. Before sending texts containing personal or sensitive information to third-party services, confirm the purpose, retention period, access rights, and whether the data is used for training; if you can't guarantee these, process locally or only send de-identified segments.

From Themes to Action Requires a Synthesis of Evidence

A publishable theme card should contain: research question, code definition, valid denominator, whether weighted or not, number covered, major population differences, representative quotes, counterexamples, correspondence with closed-ended questions or behavioral data, and recommended next steps. For the example above, the recommendation isn't a generic “improve customer service.” Instead, it would be to first examine tickets where the promised time has passed and no human follow-up occurred, distinguishing between queuing, handoffs, and lack of authorization. Then, define “faster initial response” and “clearer solution” as two testable changes, and in subsequent experiments, observe resolution time and repeat contact rates.

If your team lacks stable variables, missing value rules, and versioning, start by establishing a survey data codebook and data versioning and lineage records. For comparisons involving proportions across groups, refer to the cross-tabulation analysis guide. For final delivery, follow the actionable insight report guide to combine frequency, context, and decision thresholds. After completing these steps, open-ended findings will include both the human voice and auditable boundaries.

References for Open Text Analysis

Weiterlesen

Diese Inhalte könnten Sie ebenfalls interessieren.

Zurück zur Wissensdatenbank

Ihre Frage wurde nicht beantwortet?

Teilen Sie Ihrem Berater Ihre Forschungsziele mit – wir erarbeiten gemeinsam die nächsten Schritte.