03 · What You Need to Know
Generalizability Is Always Generalizability to Somewhere
Start by naming the target population
A statement such as “this study has limited generalizability” is incomplete unless you specify the population to which you are trying to generalize. A study can be highly informative for one population and provide much weaker evidence for another.
Suppose a randomized trial recruits adults aged 18–40 with uncomplicated disease. Its findings may be directly relevant to similar adults receiving comparable treatment. Applying the same estimate to adults over 80 with multiple comorbidities is a different inferential task.
Modern work on generalizability and transportability treats specification of the target population as a fundamental starting point. Without that target, asking whether a result is “generalizable” has no sufficiently precise meaning.
Study population
The population represented by the observations and design from which the study result is obtained.
Target population
The population to which you want the relevant estimate or conclusion to apply.
Your first task is therefore to identify whether the study sample is appropriate for the population behind the claim. Only after defining that relationship can you decide whether the inference should stop.
Different claims require different kinds of generalization
Do not ask whether an entire study generalizes as though generalizability were a single property. Identify what, specifically, is being transported.
| Claim being generalized |
What could prevent generalization? |
| Prevalence or proportion |
Different population composition, exposure patterns, selection, or determinants of the outcome |
| Population mean |
Differences in characteristics associated with the measured quantity |
| Association |
Different causal structures, confounding patterns, selection mechanisms, or measurement processes |
| Causal effect |
Different distributions of effect modifiers, versions of treatment, baseline conditions, or implementation |
| Measurement property |
Different languages, interpretations, response processes, or construct meanings |
| Implementation outcome |
Different resources, infrastructure, institutions, staffing, policies, or feasibility constraints |
A mechanism may plausibly operate in another population even when the exact numerical estimate does not transfer. Conversely, a similar demographic profile does not guarantee that an intervention will perform similarly when its implementation environment changes.
Stop when the target population contains people the study systematically excluded and the difference matters
Eligibility criteria can deliberately remove older adults, children, people with comorbidities, people taking certain medications, non-speakers of a study language, participants without particular technologies, or other groups. These restrictions may be scientifically or ethically justified.
But justified exclusion does not create evidence about the people excluded. If the excluded characteristic could plausibly alter the outcome, intervention effect, harms, feasibility, or measurement process, broader inference requires additional support.
When eligibility criteria create a population substantially narrower than the conclusion, narrowing the conclusion may be more defensible than simply acknowledging the issue in a limitations paragraph.
Stop when important target groups had little realistic opportunity to enter the study
A study can technically allow a group while effectively excluding it through recruitment. Online-only participation may miss people with limited connectivity. Recruitment through specialist clinics may miss untreated or differently treated patients. A study conducted in one language may make participation unrealistic for part of the intended population.
If important groups contribute little or no evidence and there is a plausible reason the result could differ for them, a broad population claim becomes difficult to defend.
Stop when selection into the analyzed sample could materially alter the result
Generalization becomes especially difficult when the final sample is not merely different from the target population but is selected according to characteristics related to the outcome or effect of interest.
Low participation, self-selection, complete-case analysis, and other selection processes can leave researchers observing a subset whose results do not straightforwardly describe the target population. The problem cannot be diagnosed from sample size alone.
When there is a credible mechanism through which selection could substantially distort the study's findings, the inferential boundary may need to become narrower unless appropriate design or analytical methods address the problem.
Stop when attrition leaves a different population from the one that started
Longitudinal studies introduce another difficulty. The baseline sample may initially provide a reasonable connection to the target population, but selective dropout can gradually change who remains observed.
If participants with poor outcomes, adverse effects, low engagement, severe disease, or other relevant characteristics are disproportionately lost, conclusions based on the remaining participants may not describe the original cohort well, much less a broader population.
Before generalizing longitudinal findings, determine whether attrition fundamentally changed the final sample.
Stop when the setting is part of what produced the effect
Some findings depend strongly on where they were produced. An intervention may succeed because a hospital has specialist expertise, a university has unusually strong technology infrastructure, or an organization provides implementation support unavailable elsewhere.
In such cases, the intervention cannot be separated cleanly from its environment. Applying the result elsewhere requires evidence that the relevant conditions can be reproduced or that the effect persists when those conditions change.
This is why evidence from one institution requires attention to institutional context. A large number of participants from one site does not create variation across sites.
Stop before carrying a country-specific estimate across contexts without examining what country represents
Country can proxy differences in policy, healthcare, education, economics, culture, infrastructure, language, population composition, and other contextual characteristics. Some of those differences may be irrelevant to a finding; others may be central.
The appropriate question is not whether foreign evidence is inherently non-generalizable. It is whether the contextual characteristics capable of modifying the result differ between the studied and target settings.
Before extending a result internationally, examine whether the finding has a defensible basis for applying across countries.
Stop when the construct itself may not mean the same thing
Generalization sometimes fails before statistical inference even begins. A questionnaire, diagnostic category, behavioral measure, educational construct, or psychological scale may not function equivalently in another population.
Translation alone may not establish equivalent measurement. If participants interpret questions differently, response categories function differently, or the construct has different manifestations, an apparent population difference may partly reflect measurement rather than the underlying phenomenon.
When measurement equivalence is uncertain, transporting the numerical result without qualification can be difficult to justify.
Stop when plausible effect modifiers differ substantially and you cannot account for them
For causal effects, one major concern is effect heterogeneity. Treatment or intervention effects may vary according to age, baseline risk, disease severity, prior knowledge, socioeconomic circumstances, implementation conditions, or other characteristics.
If those effect modifiers are distributed differently in the target population, the average effect observed in the study population may differ from the average effect that would occur in the target population. Generalizability and transportability methods can sometimes address these differences when the relevant variables are observed and the required assumptions are defensible.
If important modifiers are unmeasured, absent from the study, or have little overlap between populations, the broader estimate may rely heavily on assumptions rather than direct evidence.
Statistical adjustment can extend evidence, but it cannot manufacture unsupported populations
Weighting, standardization, outcome modeling, inverse-odds approaches, and related methods can sometimes estimate effects for target populations that differ from the original study sample. The literature on generalizability and transportability provides formal frameworks for doing so.
These methods require information about both study and target populations and assumptions about variables related to selection and effect heterogeneity. They do not license unrestricted extrapolation to populations with characteristics unsupported by the observed data.
In particular, poor overlap is a warning sign. If certain types of people occur in the target population but have essentially no counterparts in the study, statistical models may be asked to extrapolate beyond the empirical support available.
External evidence can justify going beyond the original sample
You do not have to demand that every study independently contain every population to which its findings might eventually apply. Scientific knowledge accumulates across studies.
Replication, systematic reviews, multisite research, studies in complementary populations, mechanistic evidence, and formal transportability analyses can provide the missing bridge. A single study may be narrow while the broader evidence base supports a wider conclusion.
This distinction is important. The correct statement may be “this study alone does not establish the broader claim,” not “the broader claim must be false.”
Refusing to generalize is not the same as rejecting the study
A finding can be valid, important, and useful within a restricted population. Narrow scope is not synonymous with weak research.
Indeed, one of the most scientifically responsible responses to uncertain external validity is simply to describe the result accurately: “among the participants studied,” “within these institutions,” “among eligible patients,” or “under these implementation conditions.”
STROBE explicitly asks authors of observational studies to discuss generalizability and to interpret findings cautiously in light of study objectives, limitations, and other evidence. Calibrating the conclusion to the evidence is therefore part of interpretation, not an embarrassing footnote appended after the interesting claims have already escaped.
Watch Out
Do not interpret “we cannot justify generalizing this result” as “we know the result will not hold elsewhere.” The first statement identifies insufficient evidence for the broader inference. The second makes a new empirical claim that also requires evidence.