01 · The Question
Are the People Who Finish the Study Still Comparable to Those Who Started It?
A longitudinal study can begin with an impressive cohort and end with something quite different. Participants move, withdraw consent, stop responding, experience adverse events, become too ill to continue, lose interest, discontinue treatment, or simply become unreachable.
If these losses occur randomly with respect to the research question, the main consequence may be reduced sample size and precision. If particular kinds of participants are disproportionately lost, however, the final analytical sample can differ systematically from the original sample.
The central issue is therefore not simply whether attrition was “high.” It is whether the process that determined who remained in the study changed the evidence in a way that threatens the estimate or conclusion.
03 · What You Need to Know
Attrition Is a Selection Process That Happens After a Study Begins
Attrition changes who contributes follow-up information
At baseline, researchers observe an initial cohort. At later time points, only some members of that cohort may still provide measurements. The resulting final sample is therefore selected partly by the process that determines continued participation.
This matters because longitudinal inference often concerns change, incidence, later outcomes, or treatment effects over time. Participants without later measurements cannot contribute the same outcome information as participants who remain observed.
STROBE reporting guidance for cohort studies therefore asks authors to report participant numbers at each stage, reasons for nonparticipation, methods of follow-up, how missing data and loss to follow-up were addressed, and relevant sensitivity analyses.
High attrition and attrition bias are not synonymous
A large percentage of missing follow-up observations is concerning because it removes information and creates greater scope for the observed sample to differ from the original cohort. But the attrition rate alone does not establish the magnitude or direction of bias.
If people are lost for reasons essentially unrelated to the quantities being estimated, the primary effect may be lower precision. If loss depends systematically on treatment, exposure, prognosis, outcome, or their causes, the observed estimate can become distorted.
Attrition
Loss of participants or their measurements after entry into a longitudinal or repeated-measures study.
Attrition bias
Systematic distortion that arises when the process determining who remains observed changes the estimate of interest.
As with low participation at initial recruitment, the percentage provides a warning signal. The mechanism determines what that warning means.
Compare baseline characteristics of retained and lost participants, but do not stop there
A common diagnostic is to compare participants who complete follow-up with those who do not. Differences in age, baseline severity, socioeconomic circumstances, treatment assignment, previous outcomes, or other measured characteristics may indicate selective retention.
These comparisons can be informative, especially when the variables are related to the primary outcome. However, absence of statistically significant baseline differences does not prove that attrition is harmless. Tests may have limited power, relevant variables may be unmeasured, and loss can depend on post-baseline information that baseline tables do not capture.
Use baseline comparisons as evidence about attrition, not as a pass-or-fail test.
Why participants leave can be more informative than how many leave
Consider a trial comparing two treatments. Some participants move away and can no longer attend follow-up visits. Others stop participating because the treatment causes intolerable adverse effects. Still others withdraw because their symptoms worsen and they seek alternative care.
These reasons are not inferentially equivalent. Attrition related to adverse effects or worsening outcomes is directly connected to the treatment's consequences and may make the retained participants look healthier or more tolerant of treatment than the original group.
Whenever reasons for withdrawal are available, examine their distribution across study groups and their relationship to the outcome.
Differential attrition can undermine comparisons between groups
Attrition becomes particularly concerning when the amount or reasons for loss differ between treatment, exposure, or comparison groups.
Suppose 8% of participants leave the control group but 35% leave the intervention group, with many intervention participants withdrawing because the program is burdensome. An analysis based only on completers could compare a relatively broad control group with a highly selected subset of intervention participants who tolerated and adhered to the program.
Even equal attrition rates across groups do not guarantee absence of bias, however. Twenty percent could leave both groups for very different outcome-related reasons.
Attrition can make the remaining sample look artificially successful
This problem is easy to imagine in behavioral, educational, and clinical interventions. Participants who struggle with an intervention may be especially likely to discontinue it. If researchers analyze only those who finish, the remaining group can become enriched with people for whom the intervention was feasible, tolerable, or effective.
The resulting estimate may then answer a subtly different question: how did the intervention perform among people who successfully remained in the study? That is not necessarily the same as how the intervention would perform among everyone who started or would be offered it in practice.
Complete-case analysis can intensify the selection
Researchers sometimes exclude every participant missing any variable needed for a model. This complete-case approach can reduce the analyzed sample far beyond the visible dropout count.
Whether complete-case analysis is unbiased depends on assumptions about why data are missing and the analytical structure. Automatically discarding incomplete observations can create or amplify selection when having complete data is related to variables relevant to the estimate.
When the final analytical N is substantially smaller than the baseline N, determine exactly where those observations went.
Missing-data terminology describes assumptions, not convenient labels
Missing-data methods often distinguish among missing completely at random, missing at random, and missing not at random. These terms have technical meanings concerning how missingness relates to observed and unobserved data.
Calling data “missing at random” does not make the assumption true. Researchers should explain why their analytical strategy is plausible given the study design and available information.
Methods such as multiple imputation, likelihood-based analyses, inverse probability weighting, and other approaches can be preferable to simplistic complete-case analysis under appropriate assumptions. None can recover information without assumptions about the missingness process.
Sensitivity analyses matter when the missingness mechanism is uncertain
The people who leave a study are often precisely the people for whom researchers cannot observe the final outcome. That makes some assumptions about their unobserved outcomes impossible to verify directly from the dataset.
Sensitivity analyses can ask how conclusions would change under alternative plausible assumptions about those missing outcomes. If a modest departure from the primary missing-data assumptions reverses the conclusion, the study's result is more fragile than a single primary analysis might suggest.
Attrition can threaten both internal and external validity
In randomized trials, differential post-randomization loss can compromise the comparability created by random assignment in the observed data. In observational longitudinal studies, selective retention can alter exposure-outcome relationships. In either design, the retained sample may also become less similar to the population to which researchers hope to generalize.
Attrition can therefore be both an internal selection problem and an external-validity problem, depending on the mechanism and target inference.
This is part of the broader question of when selection processes materially distort study findings.
A large baseline sample does not protect against selective loss
Beginning with 20,000 participants may leave a numerically impressive sample even after substantial attrition. But if the 12,000 remaining participants are systematically selected according to factors related to the outcome, the problem is not solved by the fact that 12,000 is still a large number.
As with large selectively obtained samples more generally, sample size can provide precision without eliminating systematic selection.
Watch Out
Do not judge attrition solely by a universal percentage cutoff. The same attrition rate can have very different consequences depending on who leaves, why they leave, when they leave, and how those mechanisms relate to treatment, exposure, outcomes, and the analysis.
06 · What This Means for You
Follow the Sample From Baseline to the Final Analysis
When evaluating a longitudinal study, treat participant flow as part of the evidence rather than administrative detail. Record how many participants began, how many contributed data at each major follow-up, and how many entered the final analysis.
A simple decision framework
If attrition is small and appears unrelated to variables relevant to the outcome
The main consequence may be reduced precision, although the missingness assumptions should still be considered.
If attrition is substantial and retained participants differ from those lost
Examine whether those differences plausibly affect the outcome or treatment effect and whether the analysis addresses them.
If attrition differs substantially between study groups
Investigate both the magnitude and reasons for differential loss before accepting a comparison based on observed completers.
If withdrawal is related to adverse effects, worsening outcomes, or intervention burden
Treat completer-only results particularly cautiously because remaining participants may represent a selected group that tolerated or benefited from the intervention.
If missing outcomes are substantial
Inspect the missing-data method and sensitivity analyses rather than assuming that software-generated adjustments have resolved the problem.
Finally, compare the population represented at baseline with the people represented by the final analysis. If important groups have disproportionately disappeared, the study may also face the broader problem that some relevant populations are no longer meaningfully represented in the evidence.
A useful appraisal does not simply report that “30% were lost to follow-up.” It explains what that loss may have done to the comparison or estimate. The percentage gets you to the methodological question; it does not answer it.
07 · A Quick Checklist
How to Evaluate Attrition and the Final Sample
When participants are lost during follow-up, check:
Record participant numbers at baseline, each important follow-up, and final analysis.
Examine when participants were lost and the reported reasons for withdrawal or missing follow-up.
Compare attrition rates across treatment, exposure, or other analytically important groups.
Compare retained participants with those lost on relevant observed baseline and intermediate characteristics where possible.
Ask whether leaving the study could depend on treatment effects, adverse events, worsening outcomes, prognosis, exposure, or related variables.
Determine whether the primary analysis uses all appropriate available information or only complete cases.
Inspect the assumptions underlying imputation, weighting, likelihood-based methods, or other missing-data approaches.
Look for sensitivity analyses showing whether plausible alternative assumptions about missing outcomes materially change the result.
Check whether the conclusions describe the original study population more broadly than the retained evidence supports.