Of 400 surveyed customers, 232 said they would renew, and the report states, "Renewal intention: 58%." This point estimate looks precise to the single digit but does not tell the reader how much the result might vary if another set of customers were sampled and surveyed using the same method, whether 58% is sufficient evidence to conclude it exceeds 50%, or whether the true level could be below the 60% expansion threshold.
A confidence interval addresses only the "estimation variability" part. It combines the point estimate, standard error, and a pre-selected coverage level to indicate the precision of a statistical procedure under a specified design and model. It does not automatically account for missing populations, leading questions, fraudulent responses, or voluntary participation bias. An interval can be narrow while consistently surrounding the wrong answer.
The 95% Refers to the Procedure, Not a Probability Assigned to This True Value
Under the frequentist framework, the population parameter is treated as fixed but unknown. If one were to repeatedly sample from the same population using the same design and compute a 95% confidence interval each time, approximately 95% of those intervals would, in the long run, contain the parameter. For the already calculated interval of 53.1% to 62.7%, the parameter is either inside or outside; it cannot be strictly interpreted as "there is a 95% probability that the true renewal intention falls here." The NIST definition of the confidence coefficient is based on this repeated-sampling coverage.
It also does not imply that 95% of respondents' answers fall within the interval, that there is a 95% probability the next sample's point estimate will fall within the interval, or that 95% of individuals will renew. If Bayesian credible intervals are used, a corresponding posterior probability interpretation is possible with explicit priors and models; but one cannot simply relabel a frequentist interval and use that language.
Calculating the Interval for 232/400: Report the Method Name Too
The point estimate is 232 divided by 400, or 58.0%. If we temporarily assume these 400 individuals come from a simple random sample and their responses are independent, and we use a 95% Wilson binomial proportion interval, the result is approximately:
Proportion with renewal intention: 58.0%
95% Wilson Confidence Interval: 53.1%–62.7%
Unweighted sample size: n = 400This interval supports the judgment that the population proportion exceeds 50% under the stated sampling and calculation assumptions because the lower limit exceeds 50%; however, it does not confirm the 60% business threshold is met, as the interval contains values below 60%. The point estimate of 58% itself is also below the threshold. If the action rule requires "sufficient evidence for at least 60%," the current evidence is insufficient; if the rule only requires the interval to exclude 50%, the conclusion may differ. Statistical intervals must be paired with pre-specified action rules.
Proportion intervals are not always symmetric "estimate plus or minus 1.96 standard errors." When the sample is small, success or failure counts are few, or the proportion is near 0 or 1, the simple Wald normal approximation can have poor coverage or even cross 0% or 100%. The NIST binomial proportion interval documentation provides Wilson and exact methods; in practice, one may also choose Agresti-Coull, Jeffreys, or other methods depending on the analysis objective. The key is not to list all names, but to pre-specify a method, use an appropriate implementation, and report the method in the write-up.
Same 58%, Different Sources Imply Different Commitments
| Data Source | How the Interval Should Be Calculated or Expressed | What Cannot Be Claimed |
|---|---|---|
| Simple random sample of customers | Appropriate proportion intervals for sampling variability, handling finite population conditions | Cannot cover nonresponse, question bias, or data processing errors |
| Stratified, clustered sampling with unequal probabilities | Use weights, stratification, clustering, and design-based variance | Cannot treat 400 rows as 400 independent, equally weighted observations |
| Repeated measurements from the same customer | Calculate intervals for change scores or paired structures | Cannot apply standard errors from two independent samples |
| Open link on social media or voluntary panel | If reporting an interval, describe the repeated sampling process, model, assumptions, and validation | Cannot claim traditional probability sampling margins of error based on sample size alone |
Suppose the 400 responses actually come from 40 stores, with customers from the same store being more similar. Simple random formulas will treat correlated responses as new, independent information. If a hypothetical design effect were 1.8, the standard error would be multiplied by its square root, expanding the ordinary approximate margin of error from about ±4.8 percentage points to about ±6.5 percentage points. This number only illustrates how a design effect changes precision; a real project must obtain this estimate from its sampling design and appropriate variance estimation, not uniformly apply 1.8.
The CDC's complex survey variance tutorial notes that weights affect parameter estimates, and stratification, clustering, and unequal weights affect standard errors, test statistics, and intervals. Plugging weighted "representative counts" into a generic online calculator can produce absurdly narrow intervals. Keep both the unweighted sample size, weight variables, and sampling design fields.
Precision and Bias Are Two Axes, Not Offsetting Forces
Survey results can be evaluated on a two-dimensional grid:
| Weak Evidence of Bias | Strong or Unknown Bias Risk | |
|---|---|---|
| Narrow interval | Estimate may be both stable and close to the target population; can make threshold-based decisions and disclose limitations | Result may be precisely wrong; address coverage, response, or measurement issues first |
| Wide interval | Direction may not be wrong, but information is insufficient; increase effective independent sample or accept risk | Neither precise nor likely unbiased; do not use a point estimate to support strong conclusions |
Increasing the sample from 400 to 800, all else being equal, only reduces the standard error of a proportion to about 1 divided by the square root of 2 of its original size, not half. Using the conservative normal approximation near 50%, the 95% half-width is about 4.9 percentage points for n=400, about 3.5 for n=800, and about 2.5 only for n=1600. Adding responses mainly improves random precision; if the new respondents are still the same type of active users, self-selection bias will not disappear with a larger sample.
Pew Research Center's work on variability of online opt-in samples explicitly calls such intervals model-based margins of error: they depend on assumptions about the repeated sampling process, while systematic coverage, self-selection, or nonresponse bias can cause estimates to vary around the wrong center. Consistent with this, the AAPOR Transparency Initiative stipulates that non-probability samples should provide precision measures only when the model specification, assumption checks, and calculation methods are described.
Comparing Two Groups Requires a "Difference Interval," Not Overlap of Two Intervals
If renewal intention is 62% for new customers and 54% for existing customers, drawing separate 95% intervals for each group still cannot strictly answer whether the difference exceeds zero, because the difference between the two estimates has its own standard error, and paired or shared sampling can introduce covariance. One should directly report, "New customers minus existing customers is 8 percentage points, with a difference interval of ..." using methods appropriate for independent, paired, clustered, or weighted designs.
Similarly, a rise in one quarter and a fall in the next does not establish a trend; one must estimate the change itself and its interval, and confirm that the samples, questions, weights, and response patterns are comparable. To choose between two-proportion, paired, regression, or survey design methods, start with a statistical test decision aid to clarify the estimation target and independent units.
Overall Interval Does Not Carry Over to Each Small Subgroup
The statement "overall n=400, margin of error approximately ±5 percentage points" applies only to that specific overall proportion under a simple approximation, not to a subgroup of 45 new users in South China, nor to rare events with proportions near the boundary. Each subgroup has its own unweighted base, weight variation, clustering structure, and interval; the finer the cross-tabulation, the worse the precision tends to be.
If dozens of populations and indicators are displayed simultaneously, numerous 95% intervals also invite selective interpretation: readers may only pick out the few that exclude zero. Declare primary estimates and confidence levels in the analysis plan, use appropriate simultaneous intervals or multiplicity strategies for a set of comparisons where error control is needed, and label exploratory slices.
Bringing Interval Endpoints into the Business Scenario Makes Conclusions Actionable
For the 58.0% (53.1%–62.7%) renewal intention, do not just say "margin of error approximately ±5%." Walk through each endpoint:
- If the lower bound of 53.1% is still sufficient to cover a conservative budget, the project is robust to sampling variability;
- If 60% must be reached for expansion, the interval straddles the threshold, and the current sample does not rule out "not worth expanding";
- If the financial model is sensitive to each percentage point, plug in 53.1%, 58.0%, and 62.7% separately, rather than treating the point estimate as a fixed input;
- If nonresponse analysis shows that low-activity customers are missing, run bias scenarios separately; do not force that uncertainty into the statistical interval.
The last point is especially important. Nonresponse bias diagnostics address "might non-respondents differ?" while confidence intervals address "how much would results vary under repeated sampling with the current design?" Both affect decisions, but they are not the same type of uncertainty.
Retain These Eight Items in Reports for Reproducibility
- Point estimate, interval lower and upper limits, confidence level, and interval method;
- Numerator, unweighted denominator, and the number to whom the question truly applies;
- Target population, sampling frame, and probability or non-probability recruitment approach;
- Weights, stratification, clustering, repeated measures, and finite population handling;
- Missing data, quality exclusions, extreme weights, and effective sample size;
- Whether primary indicators, subgroups, and multiple comparisons are pre-specified;
- Risks outside the interval: coverage, nonresponse, measurement, and processing issues;
- Business thresholds and the different actions implied by each end of the interval.
When viewing overall results, cross-tabs, and trends in the survey platform, preserve the true question denominator, sample source, weighting status, and version; for complex designs or model-based intervals, export to a statistical environment that supports the appropriate variance estimation, then write the method and results back into the report. Do not let AI or a generic calculator guess the sampling design from only "58%, n=400."
An adequate conclusion can be phrased as: "Among 400 customers who were simple random sampled and provided valid responses, 232 said they would renew, an estimated proportion of 58.0% with a 95% Wilson interval of 53.1% to 62.7%. The interval reflects only random variability under this sampling model; survey response and coverage limitations are discussed separately. Since the interval contains values below 60%, the current evidence is insufficient to confirm the expansion threshold is met." This is only a few lines longer than "renewal intention 58%, margin of error ±5%," yet it clearly states what the numbers can say, cannot say, and what the next step is.

