01 · The Question
Do Ten Papers From One Dataset Represent Ten Independent Pieces of Evidence?
A large cohort, national survey, biobank, registry, administrative database, or electronic health-record system can support dozens or even hundreds of research papers. Each publication may have a different title, author team, exposure, outcome, sample size, and analytical model.
Those papers may answer genuinely different research questions. What they do not necessarily provide is an equally large number of independent participant populations.
This distinction becomes consequential in systematic reviews and meta-analyses. Methodological work on real-world and observational evidence has documented how multiple studies using the same database, overlapping analytic periods, or shared participant groups can lead to double-counting when the resulting estimates are synthesized as though they were independent.
03 · What You Need to Know
Why Publication Diversity Can Conceal Data Dependence
A Dataset Can Support Many Legitimate Research Questions
Large datasets are deliberately built or maintained so that researchers can investigate more than one question. A longitudinal cohort can support studies of education, health, employment, family relationships, and later-life outcomes. A hospital database can support analyses of different diseases, treatments, and patient subgroups.
There is nothing inherently problematic about this. Reusing well-characterized data can increase research efficiency and permit questions that would be difficult or expensive to investigate through new data collection.
The methodological problem is narrower: two papers based on the same dataset may share participants, meaning their estimates may not be statistically independent even though the publications look unrelated.
Same Dataset Does Not Automatically Mean Same Sample
Two papers can use the same database and contain no shared participants. One might examine children while another examines adults, or one might use records from 2010 through 2012 while another uses newly enrolled participants from 2020 through 2022.
Conversely, two papers with different sample sizes may overlap extensively if they use the same source population, calendar period, and similar eligibility criteria.
Shared dataset
The papers obtain information from the same underlying data resource.
Overlapping analytic sample
Some or all individual participants or records included in one analysis also contribute to another.
The first should prompt investigation of the second. It does not prove it.
Large Named Cohorts Are Often Easier to Recognize
Overlap may be relatively visible when papers explicitly name a cohort, survey, biobank, or registry. You can record the dataset name and compare waves, years, age ranges, eligibility criteria, and outcome samples.
Problems become harder when the data source is described generically, such as “electronic medical records from a large health system,” or when different papers use slightly different names for the same underlying resource.
Data provenance should therefore be extracted as a study characteristic rather than left buried in the methods section.
Observation Periods Can Reveal Potential Overlap
Consider two studies using the same national registry. Paper A includes patients diagnosed between 2015 and 2020. Paper B includes patients diagnosed between 2018 and 2023. If their other eligibility criteria are similar, participants diagnosed during 2018–2020 could potentially occur in both samples.
The extent of overlap may still be unknown. The calendar windows tell you that overlap is possible, not exactly which individuals are duplicated.
Methodological work on double-counting in evidence synthesis identifies overlapping timeframes within the same database as a recurring source of sample overlap.
Different Research Questions Do Not Guarantee Independent Participants
One paper may examine mortality after Treatment A, another hospital admission after Treatment B, and a third a risk factor for disease progression. If all three draw from the same database during overlapping periods, some individuals may contribute to several analyses.
The publications can still represent distinct scientific questions. Scientific distinctness and statistical independence are different properties.
| Feature |
Can Differ Across Papers? |
Does a Difference Prove Independent Samples? |
| Title |
Yes |
No |
| Authors |
Yes |
No |
| Research question |
Yes |
No |
| Outcome |
Yes |
No |
| Sample size |
Yes |
No |
| Analysis method |
Yes |
No |
| Underlying dataset |
May be the same |
Shared dataset raises the possibility of overlap |
| Mutually exclusive participant IDs or sampling frames |
Can differ |
Can provide strong evidence of independence when adequately documented |
One Paper Can Be Nested Inside Another
Suppose Paper A analyzes all 50,000 eligible adults in a database. Paper B examines 8,000 participants with diabetes drawn from the same eligible population and observation period.
Paper B may address an important subgroup question, but its 8,000 participants may already be contained within Paper A. If both estimates are entered into the same synthesis without accounting for their relationship, the subgroup participants can effectively contribute more than once.
This is a form of overlapping study population.
Shared Controls Can Create Less Obvious Dependence
Overlap can also occur when the focal treatment or exposure groups differ but comparison participants come from the same database. Two studies may therefore appear to compare completely different interventions while sharing part of the underlying control population.
Cochrane warns against analogous double-counting within multi-arm trials because reusing a shared comparator in multiple supposedly independent comparisons creates a unit-of-analysis problem. With routinely collected data, the same conceptual issue can arise across separately published studies when shared participants are not apparent.
Real-World Data Make the Problem Particularly Difficult
Studies based on registries, claims databases, electronic health records, and other routinely collected sources can involve very large populations and complicated eligibility algorithms. The same person may satisfy the criteria for several published analyses.
Hussein and colleagues highlighted this problem in evidence synthesis involving real-world and observational studies. Their case studies included multiple papers using UK Biobank and multiple analyses derived from the same hospital databases. The authors noted that the exact extent of overlap is often impossible to determine from published reports.
This uncertainty matters because simply adding the reported sample sizes can substantially overstate the number of unique individuals represented by the literature.
Overlapping Samples Can Produce Overconfident Meta-Analysis
Conventional meta-analytic methods generally combine effect estimates from studies under assumptions about their statistical relationships. When estimates based on shared participants are treated as independent, the correlation between them is ignored.
Methodological work on overlapping real-world populations warns that double-counting can produce spuriously high precision. The exact consequences depend on the structure and magnitude of the overlap, the effect estimates, and the analytical method.
This does not mean every pair of papers sharing a dataset must be excluded. It means the dependence needs to be recognized and handled deliberately.
Publication Count Can Exaggerate Apparent Replication
The problem is not confined to statistical pooling. Suppose a literature contains 20 papers reporting broadly similar associations, but 14 derive from the same national cohort.
A narrative summary that describes the finding as reproduced across 20 independent studies would mischaracterize the evidence base. The association may have been observed repeatedly across different analyses, but independent replication across populations is more limited.
Analytical replication within a dataset
A finding is examined repeatedly using related samples, outcomes, models, or subgroups from the same data resource.
Independent replication across populations
A finding is examined in participant populations whose observations do not derive from the same underlying sample.
Both can be informative, but they support different claims about robustness and generalizability.
The Same Dataset Can Also Produce Truly Non-Overlapping Samples
A large data resource may contain independent recruitment waves, non-overlapping geographic populations, or mutually exclusive eligibility groups. Do not collapse all papers using the same dataset into one study merely because the data source is shared.
The correct task is to reconstruct the analytic cohorts. For each publication, record the source dataset, years or waves, geographic restrictions, eligibility criteria, exposures or interventions, outcomes, and sample size. Then assess whether overlap is known, probable, possible, or absent.
Exact Overlap Is Often Unknowable
Published papers usually provide aggregate descriptions rather than participant identifiers. Even if two studies use the same database and overlapping years, you may not be able to determine exactly how many people appear in both.
Do not manufacture an overlap percentage from approximate similarities. Record the uncertainty explicitly. Depending on the review, sensitivity analyses that remove potentially overlapping studies can help assess whether conclusions depend heavily on those publications.
Watch Out
Do not treat “different paper,” “different analysis,” and “independent evidence” as synonyms. A paper can be scientifically distinct while remaining statistically dependent on other papers because the analyses reuse some of the same people.
07 · A Quick Checklist
How to Detect Shared Data Behind Apparently Independent Papers
For each potentially related publication, check:
The exact cohort, registry, survey, biobank, administrative database, electronic health-record system, or other data source used.
Data-collection waves, calendar years, recruitment periods, and follow-up windows.
Geographic areas, institutions, hospitals, schools, or other sampling locations represented.
Eligibility criteria, exclusions, subgroup restrictions, and missing-data requirements used to construct each analytic sample.
Whether one sample is explicitly described as a subset or secondary analysis of another population.
Whether studies could share comparison groups even when their focal exposures or interventions differ.
Whether participant overlap is known, probable, possible, or demonstrably absent.
Whether your synthesis method assumes independence between estimates that may share participants.
Whether conclusions about replication distinguish repeated analyses of one dataset from evidence across independent populations.