GuíaPaired samples t testPre-post measurementStatistical analysis

How to Conduct a Paired Samples t Test: From Pre-Post Differences to Attrition and Causal Boundaries

Using a step-by-step example with 80 baseline participants and 64 complete pairs, this article clarifies paired conditions, mean differences, standard errors, confidence intervals, effect sizes, hypotheses about differences, sensitivity to attrition, and alternative methods.

Actualizado recientemente 2 de septiembre de 2026 22 min de lectura

Direct conclusion: The paired samples t test answers whether the population mean of paired differences deviates from a preset value, not whether two column means look different. Measurements of the same individuals before and after training, the same person rating under two randomly ordered conditions, or the same object measured by two instruments may constitute pairs; two survey waves with equal sample sizes but different participants do not. Reports should include the number of complete pairs, mean difference, standard deviation of differences, confidence interval, and effect size, while separately noting attrition, order effects, and causal limitations due to lack of a control group.

The unit of pairing is the difference, not two columns of scores

Let Yᵢ be the post-test score and Xᵢ the pre-test score for the i-th person, with a fixed direction dᵢ=Yᵢ-Xᵢ. Each row dᵢ must come from the same object and comparable measurement conditions. If identifiers are mismatched, the same person submits multiple times, or pre/post scales differ in items or scoring direction, even the most precise test will lack interpretability. When building tables, retain the participant ID, two measurement time points, raw values, difference, completion status, version, and anomaly notes; identifiers should satisfy pairing needs without exposing unnecessary identity information in the analysis table.

The statistical efficiency of paired designs comes from within-individual comparison: a consistently high scorer serves as their own reference, and stable individual differences cancel out in the differences. The stronger the correlation, the smaller the difference fluctuation typically. However, pairing is not inherently better; it requires the same individual to undergo two measurements, making practice, memory, fatigue, temporal changes, and attrition more likely. Whether to pair must be decided at the design stage, not after seeing which test gives a significant result.

Complete calculation with 64 pre-post pairs

In a training program, 80 employees are measured on a skill scale before the program, and only 64 complete the same scale four weeks later. With “post minus pre” as the direction, among these 64 complete pairs, the pre-test mean is 61.0, the post-test mean is 64.2, so the mean difference d̄=3.2 points; the sample standard deviation of the 64 individual differences is sd=8.0 points. Note that the formula requires the standard deviation of differences, not the average of the pre-test and post-test standard deviations.

The standard error of the mean difference is sd÷√n=8.0÷√64=1.0 point. Testing whether the population mean difference is zero, t=3.2÷1.0=3.20, with degrees of freedom 64-1=63; the two-sided p-value is approximately 0.002. The critical value for 95% at 63 degrees of freedom is approximately 2.00, so the interval for the mean difference is 3.2±2.00×1.0, i.e., approximately 1.2 to 5.2 points. The correct statement is: “Complete pairs improved on average by 3.2 points, 95% CI approximately 1.2 to 5.2”; it cannot be reduced to “training effect is significant.”

The paired standardized effect size dz=mean difference ÷ standard deviation of differences = 3.2÷8.0=0.40. It helps describe change relative to within-individual fluctuation, but should not be mechanically labeled with generic “small, medium, large” thresholds. If the team pre-specified that a 5-point increase is worthwhile, the current interval spans 5, so the data are compatible with improvements below the action threshold and barely compatible with those meeting it; statistical significance does not substitute for cost-benefit decisions by the team.

First examine the difference distribution and outliers

The t test makes inferences about the mean difference, and its distributional assumptions concern dᵢ, not that the pre-test and post-test columns are each normal. Plot dot plots, histograms, or normal quantile plots of differences to check skewness, heavy tails, and a few extreme changes. With larger samples, inference on the mean is typically more robust to mild deviations; with small samples and severely skewed differences, one or two observations may dominate the conclusion. Do not first conduct a normality test and then misstate “not significant” as having proven normality; graphs, data provenance, and sensitivity to outliers are more important.

For anomalous differences, return to records to determine whether it is a reverse entry, a scale version change, a mismatch, or a real important case. Data errors are corrected per pre-specified rules; true values cannot be deleted simply because they inflate the p-value. Report both analyses with and without stated outliers, or use robust methods. If switching to the Wilcoxon signed-rank test, also understand that it relies on corresponding conditions on the difference distribution and that its target is not automatically equivalent to the mean difference.

What 16 dropouts could change

A paired t test typically uses only those with data at both time points; in this example, the sample shrinks from 80 to 64, with a complete-pair rate of 64÷80=80%. If the 16 who did not complete the post-test had a pre-test mean of 72, whereas completers had 61, attrition is clearly related to baseline; “complete pairs improved by 3.2 points” cannot be automatically generalized to the original 80, nor can missing post-tests be replaced with zeros.

Conduct transparent scenario analyses rather than pretending to know missing values. If the 16 dropouts had an average change of -2 points, the scenario mean for all 80 would be (64×3.2+16×-2)÷80=2.16 points; if they dropped by 8 points on average, the scenario mean would be only (204.8-128)÷80=0.96 points. This is not an estimate of missing data but informs decision-makers how pessimistic unobserved outcomes must be to substantially weaken conclusions. Formal handling could use mixed models or multiple imputation based on prespecified missingness mechanisms, with model assumptions reported.

Pre-post change is not causal intervention effect

Even if t=3.20, it only indicates that the mean difference for complete pairs is unlikely to be zero. Without a randomized control group, practice on the same scale, concurrent business changes, natural maturation, regression to the mean, and selective attrition can all contribute to pre-post differences. If the goal is to evaluate training effectiveness, set up a concurrent control group and compare the difference in changes between groups, rather than testing only whether the training group changed. Randomization addresses between-group comparability; a paired test addresses observational correlation; these are not the same concept.

When the same person experiences conditions A and B sequentially, randomize or balance the AB/BA order, record the washout period, and judge whether learning from the first condition carries over to the second. Interface learning, brand exposure, or knowledge acquisition often cannot be eliminated by a short interval; in such cases, parallel groups may be more credible. Simply running a paired t test on A and B scores while ignoring order and carryover effects confounds treatment differences with period effects.

Which data should not use this test

A single ordered categorical item may not be suitable for treating as a continuous mean; choose ordinal models or paired rank methods based on scale properties, distribution, and interpretation goals. For before/after “yes/no” outcomes, the key data are discordant pairs (changed from no to yes vs. yes to no), and McNemar-type methods are appropriate. With more than two time points, unequal intervals, covariates, or individuals nested within teams, repeated measures or mixed-effects models are usually more appropriate. To demonstrate equivalence of two measurements or non-inferiority of a new method over an old one, “not significant” cannot be used as evidence; instead, prespecify an equivalence margin and use the corresponding interval test.

Sample size planning should also center on differences: estimate a meaningful mean difference, standard deviation of differences, significance level, power, and attrition rate. If only standard deviations of raw scores are known and not the correlation, the required sample size for a paired design remains uncertain; use a small pilot to estimate difference variability and run scenario planning with multiple plausible values.

Publishable results template

A qualified result should state: “Eighty completed baseline, and 64 completed the four-week post-test and were successfully paired; scale and scoring rules remained unchanged. Complete pairs had an average change of +3.2 points (SD of differences 8.0, 95% CI 1.2 to 5.2, t(63)=3.20, two-sided p≈0.002, dz=0.40). The 16 dropouts had higher baseline scores; scenario analysis shows that if their average change were -2 points, the overall scenario mean would drop to 2.16 points. The study had no concurrent control, so results describe pre-post change and cannot be attributed solely to training.”

Before analysis, consult the guide to choosing statistical tests for survey data to assess variable types, use handling missing survey data to review attrition, and set action thresholds in combination with statistical significance and effect size; if the research goal is causal inference, first distinguish between exploratory, descriptive, and causal research. The paired t test is only a computational tool in the evidence chain; the study design determines how strong a conclusion it can support.

Sources for paired comparisons and reporting