01 · The Question
Are These Papers Really Based on Independent Groups of People?
You have established that two papers are distinct publications, but another question remains: did they analyze different people?
This can be surprisingly difficult to determine. Researchers may publish several analyses from the same trial, cohort, survey, registry, administrative database, school sample, or other participant pool. A later paper may use the full original sample, a subgroup, participants with complete data, or people who remained at follow-up. Two papers can therefore have different sample sizes without representing independent populations.
The distinction matters whenever your interpretation assumes independence. If the same individuals contribute data more than once, treating those observations as though they came from separate groups can give that participant pool disproportionate influence in an evidence synthesis.
03 · What You Need to Know
How to Trace Participants Across Publications
Same Study and Same Participants Are Related but Different Questions
Determining whether publications arise from the same study and determining whether they analyze the same participants are closely related tasks, but they are not interchangeable.
A single longitudinal study, for example, might enroll 1,000 participants at baseline. One paper could analyze all 1,000, another could examine 400 participants with a particular characteristic, and a later follow-up might contain 720 of the original participants. The publications belong to the same broader study, yet their analytic samples are not identical.
Conversely, separate analyses or nominally separate studies can draw from the same registry, administrative database, cohort, or other source population. The relevant question is therefore participant provenance: where did these observations actually come from?
Same study
The papers originate from the same underlying research project or investigation.
Same participants
The analytic samples contain some or all of the same individuals, regardless of how the publications are labeled.
If you first need to establish the relationship between the investigations themselves, compare the characteristics used to identify papers from the same underlying study.
Start With Study and Dataset Identifiers
A shared trial registration number, cohort name, study acronym, registry identifier, protocol, ethics identifier, grant number, or named dataset can provide strong evidence that the participant pools are related.
These identifiers do not always prove that the analytic samples are identical. Two papers using the same cohort may select different waves, age groups, geographic areas, or subsets. A shared identifier tells you where to look next.
Compare Recruitment Sites and Source Populations
Ask where the participants came from. Were they recruited from the same hospitals, schools, universities, communities, clinics, registries, panels, databases, or geographic areas?
The specificity of the match matters. Two studies both described as being conducted in the United States provide little evidence of participant overlap. Two papers drawing patients from the same three specialist clinics during the same period deserve much closer examination.
For database studies, identify the actual data source rather than stopping at a generic description such as “electronic health records.” The same named national registry, insurance database, longitudinal cohort, or institutional repository can support many publications.
Recruitment Dates Are Often Crucial
Compare the periods during which participants entered the study or dataset. Identical or substantially overlapping recruitment windows, combined with the same setting and eligibility criteria, can be a strong signal of shared participants.
Suppose one paper includes patients recruited at Hospital A from January 2020 through December 2022, while another uses patients from the same hospital from January through December 2021. The second population could be contained within the first. That does not prove overlap, but it creates a plausible mechanism for it.
Non-overlapping recruitment periods, by contrast, can sometimes establish independence even when the research team, institution, and procedures are the same.
Do Not Compare Only the Final Sample Size
The number appearing in a paper's results section may represent only the final analytic sample. Researchers may begin with a larger cohort and exclude participants because of missing data, eligibility restrictions, loss to follow-up, availability of specimens, completion of a particular instrument, or requirements of a secondary analysis.
Trace the sample backward when possible:
Source population Who could potentially have entered the analysis?
Recruited or enrolled sample How many individuals actually entered the underlying study?
Eligible sample Were additional criteria applied for this particular paper?
Analytic sample How many participants ultimately contributed to the reported analysis?
Two papers reporting 800 and 620 participants could therefore have complete overlap for those 620 people, partial overlap, or no overlap at all. The numbers alone cannot tell you which.
Compare Baseline Characteristics as a Participant Fingerprint
When explicit identifiers are unavailable, baseline characteristics can provide useful corroborating evidence. Compare age, sex or gender distribution, group sizes, diagnoses, disease severity, socioeconomic characteristics, geographic distribution, baseline outcome values, and other variables relevant to the population.
Highly similar values become more informative when several distinctive characteristics match simultaneously. Still, baseline characteristics are aggregate summaries, not individual identifiers. Similar populations can produce similar descriptive statistics by chance or design.
Read the Methods for Subset Language
Authors often reveal the relationship indirectly. Look for statements indicating that the analysis used a subset, subsample, secondary analysis, nested sample, follow-up cohort, respondents with complete data, biomarker subsample, or participants from selected sites.
Phrases such as “participants were drawn from,” “among participants enrolled in,” or “we analyzed data from” can be more informative than the title or abstract. Follow any citation attached to the description of the parent study.
Recognize Full, Partial, and Uncertain Overlap
Participant overlap is not an all-or-nothing phenomenon.
| Relationship |
What It Means |
Example |
| Complete overlap |
Essentially the same participant sample is used in both papers. |
Two outcome papers analyze the same 300 randomized participants. |
| Nested sample |
One paper's participants are a subset of another paper's sample. |
A biomarker analysis uses 180 participants from a 600-person trial. |
| Partial overlap |
The samples share some participants but each also contains participants absent from the other. |
Two analyses use overlapping recruitment years from the same longitudinal cohort. |
| No overlap |
The samples are drawn independently despite other similarities. |
The same research group repeats a study with a newly recruited cohort. |
| Uncertain overlap |
The publications do not provide enough information to determine the relationship reliably. |
Two papers use the same database and overlapping dates but do not report sufficient cohort-construction details. |
The distinction between these relationships becomes particularly important when evaluating an overlapping study population.
Shared Databases Require Particular Caution
Large cohorts, registries, electronic health records, claims databases, and public datasets can generate many publications. Papers based on the same dataset should not automatically be considered duplicate studies, because researchers may ask different questions and construct different samples. Yet neither should they automatically be treated as independent populations.
Published methodological work on evidence synthesis has highlighted participant double-counting when multiple studies use the same database or overlapping analytic periods. The practical problem is especially difficult when publications provide insufficient detail to reconstruct the exact overlap.
When the same named dataset appears repeatedly, record the dataset, sampling frame, dates, inclusion criteria, exclusions, geographic restrictions, and analytic sample for each paper. This also helps identify when one dataset produces many papers that appear to be independent evidence.
Sometimes You Cannot Determine the Exact Number of Shared Participants
Published reports usually provide aggregate information. Unless authors explicitly describe the overlap or provide identifiers that permit linkage, you may be able to conclude only that overlap is probable or possible rather than calculate exactly how many individuals are shared.
Do not convert suspicion into an invented percentage. Record what can be established, what remains uncertain, and why. If the distinction materially affects the synthesis, additional reports, protocols, registries, supplementary materials, or correspondence with investigators may clarify the relationship.
Watch Out
Different sample sizes do not establish independent samples. A smaller paper may simply analyze a subset of participants from a larger study, and a later paper may contain fewer participants because of attrition or missing data.
07 · A Quick Checklist
What to Check Before Calling Two Samples Independent
For potentially related papers, check:
The parent study, cohort, trial, registry, database, or dataset from which participants were obtained.
Study identifiers, registration numbers, cohort names, protocols, and other identifiers connecting the samples.
Recruitment institutions, geographic areas, and other sampling locations.
Recruitment, enrollment, observation, and follow-up dates.
Eligibility criteria and any additional restrictions imposed for the particular analysis.
Original cohort size, final analytic sample, group allocation, attrition, and missing-data exclusions.
Baseline characteristics that may corroborate or challenge a suspected match.
Language indicating a subset, secondary analysis, follow-up cohort, nested sample, or reuse of an existing dataset.
Whether the evidence supports complete overlap, partial overlap, no overlap, or only an uncertain relationship.