Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can the Same Dataset Produce Many Papers That Look Like Independent Evidence?

Large cohorts, registries, surveys, and administrative databases can generate many papers without generating equally many independent participant populations. Learn how to distinguish legitimate dataset reuse from independent replication.

283
Multiple Papers From the Same Dataset Guide 283 of 899
01 · The Question

Do Ten Papers From One Dataset Represent Ten Independent Pieces of Evidence?

A large cohort, national survey, biobank, registry, administrative database, or electronic health-record system can support dozens or even hundreds of research papers. Each publication may have a different title, author team, exposure, outcome, sample size, and analytical model.

Those papers may answer genuinely different research questions. What they do not necessarily provide is an equally large number of independent participant populations.

This distinction becomes consequential in systematic reviews and meta-analyses. Methodological work on real-world and observational evidence has documented how multiple studies using the same database, overlapping analytic periods, or shared participant groups can lead to double-counting when the resulting estimates are synthesized as though they were independent.

02 · The Short Answer

Yes, One Dataset Can Generate Many Distinct Papers Without Independent Samples

In Brief

The same dataset can generate many legitimate papers, but those publications should not automatically be interpreted as independent evidence because their analytic samples may overlap partly, completely, or in ways that cannot be determined from the reports.

Trace the underlying data source, sampling frame, observation period, eligibility criteria, geographic coverage, and analytic cohort for each paper. Dataset reuse is not itself a flaw; the problem arises when shared data provenance or overlapping participants are ignored when independence matters to the synthesis or interpretation.

03 · What You Need to Know

Why Publication Diversity Can Conceal Data Dependence

A Dataset Can Support Many Legitimate Research Questions

Large datasets are deliberately built or maintained so that researchers can investigate more than one question. A longitudinal cohort can support studies of education, health, employment, family relationships, and later-life outcomes. A hospital database can support analyses of different diseases, treatments, and patient subgroups.

There is nothing inherently problematic about this. Reusing well-characterized data can increase research efficiency and permit questions that would be difficult or expensive to investigate through new data collection.

The methodological problem is narrower: two papers based on the same dataset may share participants, meaning their estimates may not be statistically independent even though the publications look unrelated.

Same Dataset Does Not Automatically Mean Same Sample

Two papers can use the same database and contain no shared participants. One might examine children while another examines adults, or one might use records from 2010 through 2012 while another uses newly enrolled participants from 2020 through 2022.

Conversely, two papers with different sample sizes may overlap extensively if they use the same source population, calendar period, and similar eligibility criteria.

Shared dataset The papers obtain information from the same underlying data resource.
Overlapping analytic sample Some or all individual participants or records included in one analysis also contribute to another.

The first should prompt investigation of the second. It does not prove it.

Large Named Cohorts Are Often Easier to Recognize

Overlap may be relatively visible when papers explicitly name a cohort, survey, biobank, or registry. You can record the dataset name and compare waves, years, age ranges, eligibility criteria, and outcome samples.

Problems become harder when the data source is described generically, such as “electronic medical records from a large health system,” or when different papers use slightly different names for the same underlying resource.

Data provenance should therefore be extracted as a study characteristic rather than left buried in the methods section.

Observation Periods Can Reveal Potential Overlap

Consider two studies using the same national registry. Paper A includes patients diagnosed between 2015 and 2020. Paper B includes patients diagnosed between 2018 and 2023. If their other eligibility criteria are similar, participants diagnosed during 2018–2020 could potentially occur in both samples.

The extent of overlap may still be unknown. The calendar windows tell you that overlap is possible, not exactly which individuals are duplicated.

Methodological work on double-counting in evidence synthesis identifies overlapping timeframes within the same database as a recurring source of sample overlap.

Different Research Questions Do Not Guarantee Independent Participants

One paper may examine mortality after Treatment A, another hospital admission after Treatment B, and a third a risk factor for disease progression. If all three draw from the same database during overlapping periods, some individuals may contribute to several analyses.

The publications can still represent distinct scientific questions. Scientific distinctness and statistical independence are different properties.

Feature Can Differ Across Papers? Does a Difference Prove Independent Samples?
Title Yes No
Authors Yes No
Research question Yes No
Outcome Yes No
Sample size Yes No
Analysis method Yes No
Underlying dataset May be the same Shared dataset raises the possibility of overlap
Mutually exclusive participant IDs or sampling frames Can differ Can provide strong evidence of independence when adequately documented

One Paper Can Be Nested Inside Another

Suppose Paper A analyzes all 50,000 eligible adults in a database. Paper B examines 8,000 participants with diabetes drawn from the same eligible population and observation period.

Paper B may address an important subgroup question, but its 8,000 participants may already be contained within Paper A. If both estimates are entered into the same synthesis without accounting for their relationship, the subgroup participants can effectively contribute more than once.

This is a form of overlapping study population.

Shared Controls Can Create Less Obvious Dependence

Overlap can also occur when the focal treatment or exposure groups differ but comparison participants come from the same database. Two studies may therefore appear to compare completely different interventions while sharing part of the underlying control population.

Cochrane warns against analogous double-counting within multi-arm trials because reusing a shared comparator in multiple supposedly independent comparisons creates a unit-of-analysis problem. With routinely collected data, the same conceptual issue can arise across separately published studies when shared participants are not apparent.

Real-World Data Make the Problem Particularly Difficult

Studies based on registries, claims databases, electronic health records, and other routinely collected sources can involve very large populations and complicated eligibility algorithms. The same person may satisfy the criteria for several published analyses.

Hussein and colleagues highlighted this problem in evidence synthesis involving real-world and observational studies. Their case studies included multiple papers using UK Biobank and multiple analyses derived from the same hospital databases. The authors noted that the exact extent of overlap is often impossible to determine from published reports.

This uncertainty matters because simply adding the reported sample sizes can substantially overstate the number of unique individuals represented by the literature.

Overlapping Samples Can Produce Overconfident Meta-Analysis

Conventional meta-analytic methods generally combine effect estimates from studies under assumptions about their statistical relationships. When estimates based on shared participants are treated as independent, the correlation between them is ignored.

Methodological work on overlapping real-world populations warns that double-counting can produce spuriously high precision. The exact consequences depend on the structure and magnitude of the overlap, the effect estimates, and the analytical method.

This does not mean every pair of papers sharing a dataset must be excluded. It means the dependence needs to be recognized and handled deliberately.

Publication Count Can Exaggerate Apparent Replication

The problem is not confined to statistical pooling. Suppose a literature contains 20 papers reporting broadly similar associations, but 14 derive from the same national cohort.

A narrative summary that describes the finding as reproduced across 20 independent studies would mischaracterize the evidence base. The association may have been observed repeatedly across different analyses, but independent replication across populations is more limited.

Analytical replication within a dataset A finding is examined repeatedly using related samples, outcomes, models, or subgroups from the same data resource.
Independent replication across populations A finding is examined in participant populations whose observations do not derive from the same underlying sample.

Both can be informative, but they support different claims about robustness and generalizability.

The Same Dataset Can Also Produce Truly Non-Overlapping Samples

A large data resource may contain independent recruitment waves, non-overlapping geographic populations, or mutually exclusive eligibility groups. Do not collapse all papers using the same dataset into one study merely because the data source is shared.

The correct task is to reconstruct the analytic cohorts. For each publication, record the source dataset, years or waves, geographic restrictions, eligibility criteria, exposures or interventions, outcomes, and sample size. Then assess whether overlap is known, probable, possible, or absent.

Exact Overlap Is Often Unknowable

Published papers usually provide aggregate descriptions rather than participant identifiers. Even if two studies use the same database and overlapping years, you may not be able to determine exactly how many people appear in both.

Do not manufacture an overlap percentage from approximate similarities. Record the uncertainty explicitly. Depending on the review, sensitivity analyses that remove potentially overlapping studies can help assess whether conclusions depend heavily on those publications.

Watch Out

Do not treat “different paper,” “different analysis,” and “independent evidence” as synonyms. A paper can be scientifically distinct while remaining statistically dependent on other papers because the analyses reuse some of the same people.

04 · A Practical Example

How Four Different Papers Can Trace Back to One Dataset

Hypothetical Example

Four Studies Using the National Learning Cohort

Imagine that your review identifies four publications examining digital learning and academic outcomes.

Paper A Uses 18,000 students from the National Learning Cohort surveyed between 2022 and 2024.
Paper B Uses 12,500 students from the same cohort and years but restricts the sample to public universities.
Paper C Uses 7,200 first-year students from the same 2022–2024 cohort.
Paper D Uses 9,000 students newly recruited into the same data infrastructure in 2026.

Papers A, B, and C may contain substantial participant overlap. Paper B and Paper C are not necessarily identical subsets of Paper A, but their shared source and observation period make independence implausible without further evidence.

Paper D presents a different situation. If the 2026 wave consists entirely of newly recruited participants who were not eligible for or present in the earlier cohort, it may provide an independent participant population despite using the same data infrastructure.

The dataset name alone therefore does not determine the answer. You need the sampling frame behind each analysis.

05 · What Researchers Often Get Wrong

Common Mistakes When One Dataset Generates Many Papers

Misconception

Every Published Paper Is an Independent Study

Publication creates a new report, not necessarily a new participant population. Papers using the same dataset can contain partially or completely overlapping samples.

Misconception

Every Paper Using the Same Dataset Is a Duplicate

Also incorrect. Different papers can ask genuinely different questions and may even use mutually exclusive samples. Dataset reuse should trigger assessment of overlap, not automatic exclusion or collapsing.

Misconception

Different Sample Sizes Mean the Samples Are Independent

Different eligibility criteria and missing-data requirements can create different-sized samples containing many of the same people. Trace the sampling frame rather than relying on participant counts.

Misconception

Different Authors Mean Different Data

Large datasets are often available to many research groups. Completely different author teams can analyze overlapping participants from the same underlying resource.

Misconception

More Papers Automatically Mean More Independent Replication

Repeated findings within one dataset can provide useful analytical corroboration, but they should not be described as equivalent to replication across independent populations.

Misconception

You Can Always Determine Exactly How Many Participants Overlap

Often you cannot. Published descriptions may establish that overlap is possible or probable without permitting exact linkage. Preserve that uncertainty rather than inventing precision.

06 · What This Means for You

Extract Data Provenance Alongside the Results

When your review includes secondary analyses, routinely collected data, registries, cohorts, or large public datasets, add data provenance to your extraction process. A paper's effect estimate is easier to interpret when you know exactly which population generated it.

A simple decision framework

If papers use different named datasets
Do not assume independence solely from the names, but investigate further only when there is a plausible connection between the underlying populations.
If papers use the same dataset but mutually exclusive sampling frames
Document the evidence supporting participant independence.
If papers use the same dataset, overlapping dates, and similar eligibility criteria
Treat participant overlap as a substantive possibility and investigate before synthesizing the estimates as independent.
If one sample is clearly nested within another
Record the nested relationship and avoid treating the two samples as wholly independent.
If the extent of overlap cannot be established
Document the uncertainty and consider sensitivity analyses or other methods appropriate to the synthesis.

The aim is not to penalize productive datasets. It is to ensure that the apparent volume of literature remains proportional to the diversity of the underlying evidence. This becomes especially important when deciding how to prevent a prolific dataset from dominating your understanding of the literature.

07 · A Quick Checklist

How to Detect Shared Data Behind Apparently Independent Papers

For each potentially related publication, check:
The exact cohort, registry, survey, biobank, administrative database, electronic health-record system, or other data source used.
Data-collection waves, calendar years, recruitment periods, and follow-up windows.
Geographic areas, institutions, hospitals, schools, or other sampling locations represented.
Eligibility criteria, exclusions, subgroup restrictions, and missing-data requirements used to construct each analytic sample.
Whether one sample is explicitly described as a subset or secondary analysis of another population.
Whether studies could share comparison groups even when their focal exposures or interventions differ.
Whether participant overlap is known, probable, possible, or demonstrably absent.
Whether your synthesis method assumes independence between estimates that may share participants.
Whether conclusions about replication distinguish repeated analyses of one dataset from evidence across independent populations.
08 · Frequently Asked Questions

Questions About Multiple Papers From the Same Dataset

Are papers using the same dataset automatically the same study?

No. They may address different questions and use different, potentially non-overlapping samples. Determine the analytic population behind each paper rather than classifying studies solely by dataset name.

Can different research teams unknowingly analyze some of the same people?

Yes. Widely accessible cohorts, registries, and routinely collected databases can be analyzed by independent research teams. Different authorship therefore does not guarantee independent participants.

Does overlap matter if the papers examine different outcomes?

It depends on how the results are being interpreted or synthesized. Different outcomes can legitimately answer different questions, but the publications still should not be described as independent participant populations when they share individuals.

Can studies from the same dataset ever be independent?

Yes. They may use mutually exclusive recruitment waves, geographic areas, age groups, or other sampling frames. Independence should be supported by the actual cohort construction rather than assumed from different titles or sample sizes.

What if I know the samples probably overlap but cannot calculate how much?

Document the uncertainty. Compare data sources, periods, locations, and eligibility criteria, and consider whether sensitivity analyses excluding potentially overlapping estimates materially change the synthesis.

Why is this a problem in meta-analysis?

Effect estimates based on overlapping participants can be statistically correlated. Treating them as independent may give repeated observations excessive weight and can produce spuriously high precision.

Does using the same dataset mean one paper should simply be deleted?

No. Both papers may answer different legitimate questions. The appropriate response depends on the degree of participant overlap and the particular synthesis being conducted, not on a blanket rule that only one publication from a dataset can be retained.

09 · The Bottom Line

Many Papers Do Not Necessarily Mean Many Independent Populations

The Bottom Line

The same dataset can generate many scientifically distinct papers while providing fewer independent participant populations, so trace the source data, sampling periods, eligibility criteria, and analytic cohorts before treating those publications as independent evidence.

Dataset reuse is not inherently a methodological problem. The problem is losing sight of data provenance. Distinguishing repeated analyses within one data resource from replication across independent populations gives a more accurate account of how broad, precise, and reproducible the evidence really is.

10 · Sources and Further Reading

Authoritative and Methodological Sources on Shared Datasets and Participant Overlap

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes