01 · The Question
What Does It Mean When Study Populations Overlap?
Imagine that your systematic review contains six papers. At first glance, you have six separate pieces of evidence. Then you discover that two papers drew patients from the same national registry during overlapping years, while another paper analyzed a subgroup from one of those populations.
You may still have six publications and potentially several distinct analyses, but you no longer have six completely independent groups of participants.
This is the problem of overlapping study populations. It matters because many forms of evidence synthesis depend, explicitly or implicitly, on understanding which observations are independent. If the same people repeatedly enter the evidence base without that dependence being recognized, a particular population can exert more influence than its actual size warrants.
02 · The Short Answer
Overlap Means Some People Contribute to More Than One Study or Analysis
In Brief
An overlapping study population occurs when two or more studies, reports, or analytic samples include some or all of the same participants, so the resulting evidence is not based on completely independent groups of people.
Overlap can arise from multiple publications of one study, subgroup and follow-up analyses, repeated use of cohorts or registries, overlapping database periods, or shared comparison groups. It matters most when the analyses are treated as independent despite sharing participants, because this can create double-counting and statistical dependence.
03 · What You Need to Know
How Overlapping Populations Arise and Why They Complicate Evidence Synthesis
Overlap Exists on a Continuum
Two samples do not have to be identical to overlap. At one extreme, two papers may analyze exactly the same participants. At the other, only a relatively small subset may be shared.
Type of Relationship
Example
What Is Shared?
Complete overlap
Two papers analyze the same 500 trial participants for different outcomes.
Essentially the entire participant sample
Nested overlap
A secondary analysis uses 200 participants from a parent cohort of 1,000.
All participants in the smaller analysis occur in the larger population
Partial overlap
Two registry studies use 2018–2022 and 2020–2024 data with similar eligibility criteria.
Potentially some, but not all, participants
Shared comparison group
Two comparisons use the same control participants against different intervention groups.
The common comparator
No overlap
Two studies recruit separate cohorts during non-overlapping periods.
No participants
The methodological consequences depend on how much overlap exists, where it occurs, what effect estimates are being combined, and how the synthesis handles dependence. Merely labeling overlap as present or absent can therefore conceal important differences.
Multiple Papers From One Study Are an Obvious Source
A single investigation can generate a primary results paper, secondary outcome paper, subgroup analysis, harms report, long-term follow-up, and other publications. Cochrane guidance explicitly treats the study rather than the individual report as the primary unit of interest in systematic reviews and requires multiple reports of the same study to be collated.
If every publication is mistakenly entered as a separate study, the same participant pool may be represented repeatedly. This is why establishing whether multiple papers represent one study or several pieces of evidence is more than a bibliographic housekeeping exercise.
Overlap Can Also Occur Across Apparently Separate Studies
The problem is broader than duplicate publication. Researchers may conduct separate analyses using the same cohort, registry, electronic health record system, administrative database, biobank, household survey, or other data resource.
Consider two observational papers using the same national health database. One examines records from 2015 through 2020 and another from 2018 through 2023. If their eligibility criteria are similar, some individuals may appear in both analytic cohorts even though the publications address different research questions and are presented as separate studies.
Methodological literature has specifically identified reuse of the same database and overlapping analytic timeframes as routes through which participants can be double-counted in evidence synthesis.
A Follow-Up Sample May Overlap Without Being Identical
Longitudinal research naturally creates related populations. Suppose 800 people enter a cohort at baseline and 600 remain at a three-year follow-up. A baseline paper and the follow-up paper do not contain identical analytic samples, but the 600 retained participants are not new independent observations.
The same issue occurs when a follow-up paper uses only part of the original sample . The changing sample size should not obscure participant continuity.
Why Double-Counting Is Different From Simply Having More Evidence
If two genuinely independent studies each enroll 500 people, together they provide observations from 1,000 distinct participants. If two papers each analyze the same 500 people, there are still only 500 distinct participants behind those analyses.
Additional independent evidence
New participants contribute information that is statistically independent of participants in the other study, subject to the study design and analysis.
Repeated evidence from overlapping participants
Some of the underlying observations contribute to more than one estimate, creating dependence that must be recognized.
This does not make the second analysis useless. It may address a different outcome, follow-up period, exposure, subgroup, or research question. The problem arises when dependence is ignored and repeated observations are treated as though they came from entirely new people.
Overlap Can Violate an Independence Assumption
Conventional meta-analytic models commonly treat study effect estimates as statistically independent. When two estimates are calculated partly or wholly from the same participants, they may be correlated.
Cochrane guidance similarly notes that multiple effects calculated from the same participants are statistically dependent and that multiplicity may arise from multiple outcomes, analyses, or reports of the same study. Appropriate handling may involve selecting an eligible effect according to prespecified rules or using analytical methods that account for dependence.
The important point is not that every instance of overlap invalidates a meta-analysis. Rather, unrecognized overlap can make an independence assumption inappropriate.
Overlap Can Give Some Populations Too Much Influence
Suppose one regional registry generates eight eligible publications while another region is represented by one study. If the eight registry analyses contain many of the same patients and all are treated as independent, the first population may effectively receive repeated representation.
This can affect weighting and precision and can also shape the substantive impression of the literature. A research area may look larger and more diverse than it really is because publication count is being mistaken for population diversity.
This is one reason to examine whether the same dataset is generating many apparently independent papers and, at a broader level, to prevent a prolific dataset from dominating your understanding of the literature .
Not Every Overlap Requires the Same Response
How overlap should be handled depends on what you are synthesizing. Two papers using the same participants may report different outcomes, in which case both may be needed to characterize the study. Two correlated estimates of the same outcome may require a choice between estimates or an analytical approach capable of accommodating dependence.
Shared controls in multi-arm studies create another form of dependence. Repeated measurements and multiple follow-up periods can do the same. The methodological response should therefore be tied to the particular estimand and synthesis rather than to a blanket rule that one of every two overlapping papers must be discarded.
Watch Out
Do not solve participant overlap simply by deleting every secondary or related publication. Cochrane guidance cautions against discarding secondary reports because they may contain valuable information about study design, conduct, outcomes, and results. The objective is to avoid treating dependent evidence as independent, not to throw away useful information.
Sometimes Overlap Can Be Suspected but Not Quantified
Real-world datasets create a particularly awkward problem. Papers may report the database and observation years without identifying which individual records are shared. You may therefore know that overlap is possible or highly plausible without knowing its exact magnitude.
That uncertainty should be documented rather than converted into an arbitrary estimate. Depending on its importance, sensitivity analyses or alternative synthesis decisions may help assess whether conclusions depend on potentially overlapping studies.
06 · What This Means for You
Map Overlap Before Deciding How to Handle It
The first task is identification, not deletion. Map the parent studies, datasets, participant pools, recruitment periods, subgroups, and follow-up waves represented by your included reports. Then determine which effect estimates relevant to a particular synthesis could share participants.
A simple decision framework
If reports use the same participants and describe the same underlying study
Link the reports and treat the study, rather than each publication, as the unit of interest.
If one analytic sample is nested within another
Record the nesting explicitly and determine whether both estimates are needed for the synthesis.
If separate studies use partly overlapping populations
Assess how the overlap affects independence for the specific estimates being synthesized.
If overlap is plausible but cannot be quantified
Document the uncertainty and consider whether alternative inclusion or sensitivity analyses are appropriate.
If the samples are demonstrably independent
Record the evidence supporting that conclusion, particularly when the papers share a dataset, research team, or setting.
Whatever approach you choose should be transparent and, where possible, prespecified. The objective is not to force every complicated literature into perfectly discrete boxes. It is to ensure that the apparent quantity of evidence is not confused with the number of independent participants behind it.
07 · A Quick Checklist
How to Check for Overlapping Study Populations
Before treating included samples as independent, check:
Whether publications arise from the same parent study, trial, cohort, registry, database, biobank, survey, or research program.
Whether recruitment or observation periods overlap.
Whether recruitment sites, geographic areas, and source populations match.
Whether eligibility and exclusion criteria could select the same individuals.
Whether one sample is explicitly described as a subset, subgroup, secondary analysis, or follow-up of another.
Whether studies share a comparator or control group.
Whether apparently separate effect estimates are calculated from some of the same participants.
Whether the exact extent of overlap is known, estimated from defensible information, or genuinely uncertain.
Whether your extraction and synthesis procedures explicitly document how overlapping populations are handled.
09 · The Bottom Line
Publication Count Is Not Participant Count
The Bottom Line
An overlapping study population exists when some or all participants contribute to more than one study, report, comparison, or effect estimate, meaning those pieces of evidence are not based on completely independent samples.
Overlap does not automatically make an analysis unusable. It means you need to identify the dependence, determine whether it matters for the synthesis you are conducting, and handle it transparently rather than allowing the same participant population to acquire unintended extra influence.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation