Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Prevent a Prolific Dataset From Dominating Your Understanding of the Literature?

A widely used cohort, registry, or database can generate many papers and disproportionately shape a literature. Learn how to map shared data provenance, distinguish publication volume from independent replication, and prevent repeated samples from receiving unintended influence.

284
Preventing Dataset Dominance Guide 284 of 899
01 · The Question

What If Much of the Literature Comes From the Same Dataset?

You identify 30 papers relevant to your review. On closer inspection, 12 use the same national cohort, another five analyze one hospital database, and the remaining 13 come from separate populations. Is it still reasonable to describe the literature as 30 broadly independent investigations?

Probably not. A productive dataset can generate many scientifically legitimate papers while repeatedly representing the same underlying population. If publication count is mistaken for population diversity, one dataset can acquire disproportionate influence over both quantitative synthesis and the narrative impression of how extensively a finding has been replicated.

The solution is not to penalize productive datasets or arbitrarily retain only one paper from each database. It is to make data provenance visible, identify participant overlap where it matters, and distinguish repeated analyses of one data resource from evidence generated across independent populations.

02 · The Short Answer

Map the Evidence by Dataset and Population, Not Just by Publication

In Brief

To prevent a prolific dataset from dominating your understanding of the literature, identify which papers derive from the same underlying data source, map their analytic populations and overlap, and interpret publication volume separately from independent population replication.

There is no universal rule requiring one paper per dataset. Some analyses may use independent samples or answer genuinely different questions. The appropriate response depends on the review question and synthesis, and may involve selecting among overlapping estimates, accounting analytically for dependence, grouping related evidence, and testing whether conclusions change when heavily reused datasets receive less representation.

03 · What You Need to Know

Why a Highly Productive Dataset Can Distort the Apparent Evidence Base

Start by Separating Papers, Analyses, Datasets, and Populations

Four quantities can easily become conflated: the number of publications, the number of analyses, the number of underlying datasets, and the number of independent participant populations. They are not interchangeable.

A national cohort might support 20 eligible papers. Those papers could contain 20 distinct analyses while drawing repeatedly from one participant pool. Another literature containing 20 papers from 20 independently recruited populations has a very different evidential structure even though the publication count is identical.

Publication volume How many papers or reports exist.
Independent population evidence How many substantively independent participant populations contribute information relevant to the question.

This distinction follows the broader principle that systematic reviews should not equate reports with studies. Cochrane emphasizes that studies rather than individual reports are the primary units of interest and that multiple reports of one study should be identified and linked. The same logic becomes more complicated when nominally separate studies reuse a common database or cohort.

A Prolific Dataset Is Not Inherently a Problem

Large cohorts, biobanks, administrative databases, registries, longitudinal surveys, and electronic health-record systems exist partly so that they can support multiple investigations. Reanalysis can answer new questions, examine different outcomes, test subgroups, extend follow-up, and improve understanding of a population.

Restricting every dataset to one publication would therefore be methodologically crude. Two papers from the same resource may analyze mutually exclusive participants, different waves, or genuinely different outcomes relevant to separate syntheses.

The concern arises when repeated use of the resource is mistaken for independent replication or when overlapping estimates enter an analysis under an assumption of independence.

First Establish How Much Dataset Reuse Exists

Data provenance should be extracted explicitly. For every included paper, record the underlying cohort, registry, survey, database, biobank, health system, or other data source whenever it can be identified.

Then record the features needed to reconstruct each analytic population:

  • recruitment or observation years;
  • study waves or releases;
  • geographic and institutional coverage;
  • age or population restrictions;
  • eligibility and exclusion criteria;
  • exposure or intervention definitions;
  • outcome requirements;
  • analytic sample size; and
  • subgroup restrictions.

This converts an apparently flat list of papers into a map of where the evidence actually came from.

Same Dataset Does Not Necessarily Mean Same Participants

A shared data source is a warning to investigate overlap, not proof that two papers contain identical people.

A national database might contain mutually exclusive geographic regions. A longitudinal cohort may recruit new waves of participants. Two papers could use non-overlapping age groups or calendar periods. In such situations, the analyses may be substantially or completely independent despite sharing the same infrastructure.

Conversely, slightly different eligibility criteria can produce two samples that overlap extensively without being identical. This is why you need to determine whether multiple papers from the same dataset actually represent independent evidence.

Map Dataset Families Before Looking at the Overall Pattern

A useful organizational strategy is to group papers by their underlying data source before interpreting how many independent sources support a conclusion.

Dataset Family Publications Population Relationship Interpretive Implication
National Cohort A 10 papers Substantial overlap across waves and subgroups Many analyses, but not 10 independent replications
Hospital Database B 4 papers Overlapping calendar periods Potential dependence requires further assessment
Survey C 3 papers Mutually exclusive survey waves May provide more independent samples if participants do not recur
Independent studies 8 papers Separately recruited populations Provide distinct population-level evidence

The point is not to assign a simplistic weight to each row. It is to prevent ten papers generated from one cohort from visually masquerading as ten unrelated populations.

Do Not Count Repeated Analyses as Independent Replication

Replication usually carries a stronger implication than repeated analysis. If several research teams analyze overlapping samples from the same cohort and obtain similar findings, that may demonstrate robustness to certain analytical choices or research teams. It does not provide the same evidence about transportability across populations as reproducing the finding in independently recruited samples.

Within-dataset corroboration Related analyses from the same data resource produce compatible findings under different questions, models, subgroups, or research teams.
Across-population replication Compatible findings arise in participant populations whose observations are substantively independent of the original dataset.

Both forms of evidence can matter. They simply answer somewhat different questions about robustness.

Dataset Dominance Can Affect Narrative Reviews Even Without Meta-Analysis

Suppose 14 of 18 papers report a positive association. That initially sounds like broad consistency. If 11 of those positive papers use overlapping samples from one cohort while three independent studies produce mixed findings, the evidential picture is more complicated.

A narrative synthesis that simply counts positive and negative papers would give the prolific cohort considerable implicit weight. Even without a formal statistical model, publication frequency can shape the reader's impression of consistency.

Report the clustering directly. Instead of saying that “14 studies independently found the association,” describe how many analyses came from each major data source and how consistently the finding appeared across genuinely distinct populations.

Dataset Dominance Can Also Affect Meta-Analysis

The problem becomes statistical when estimates based on overlapping participants are entered into a meta-analysis as though they were independent.

Hussein and colleagues describe double-counting as the inclusion of the same individuals multiple times in one evidence synthesis and identify reuse of the same database, overlapping analytical periods, and common treatment groups as mechanisms through which it can occur.

The consequences depend on the extent of overlap and the analysis, but ignoring dependence can give repeated observations inappropriate influence and may make pooled estimates appear more precise than warranted.

Simply Choosing the Largest Paper Is Not a Universal Solution

One practical response to overlap is to select a single analysis from a cluster. For example, Hussein and colleagues describe a case in which several studies used UK Biobank data for a related question; the review retained one study after considering publication status and sample size.

That illustrates one possible strategy, not a universal rule. The largest paper may use a less appropriate outcome, a different population, or an inferior time point for your review. The newest paper is not automatically preferable either.

If selection is necessary, criteria should be tied to the review question and preferably specified in advance. Relevant considerations might include population fit, outcome definition, follow-up, methodological suitability, completeness, and risk of bias rather than sample size alone.

One Dataset May Legitimately Contribute to Several Different Syntheses

A cohort could provide an eligible mental-health outcome in one paper and an academic outcome in another. If your review analyzes those outcomes separately, both reports may appropriately contribute to their respective syntheses.

The key question is not whether a dataset has already appeared somewhere in the review. It is whether the same or overlapping participants are being given repeated independent representation within the particular inference you are making.

This is why blanket rules such as “one publication per dataset” can solve one problem by creating another.

Consider Sensitivity Analyses When Dataset Dominance Could Change the Conclusion

When several eligible estimates derive from potentially overlapping populations and no single analytical solution is clearly preferable, sensitivity analysis can be informative.

For example, you might compare the primary synthesis with an analysis retaining only one appropriately selected estimate from each substantially overlapping dataset family. If the conclusion changes materially, the evidence is sensitive to how repeated dataset contributions are handled.

Hussein and colleagues demonstrated empirically that removing overlapping population papers could change pooled estimates, uncertainty, and heterogeneity in a real-world case study. That does not imply the same magnitude or direction of change in every review, but it shows that overlap need not be inconsequential.

Dataset Diversity Is Not the Same as Population Diversity

Even several technically distinct databases can represent similar populations. Conversely, one large multinational resource can contain heterogeneous populations. Counting dataset names therefore remains an imperfect proxy for evidential diversity.

Consider the substantive characteristics of the populations as well: geography, healthcare system, socioeconomic context, age, ethnicity where relevant and appropriately reported, institutional setting, period of data collection, and other factors that affect the applicability of the evidence.

The goal is not to maximize the number of dataset names. It is to understand where the evidence comes from and how broadly the finding has actually been examined.

Make Dataset Provenance Visible in Your Results

If one or two data resources generate a substantial fraction of the included literature, readers should be able to see that. Depending on the review, this information may appear in study-characteristics tables, evidence maps, narrative synthesis, supplementary tables, or sensitivity analyses.

Transparency is especially important because publication titles and author lists may otherwise conceal the shared source. A reader seeing ten citations should not have to perform bibliographic archaeology to discover that eight arose from one cohort.

Watch Out

Do not “correct” dataset dominance by automatically deleting all but one paper from every shared data source. The relevant unit depends on the review question, outcome, participant overlap, and synthesis. First map the dependence; then decide what methodological response it requires.

04 · A Practical Example

When 20 Papers Represent Far Fewer Independent Populations

Hypothetical Example

A Literature Dominated by the National Student Cohort

Imagine that your systematic review identifies 20 papers examining digital technology use and student well-being.

Initial impression Fourteen of the 20 papers report an association. Read at publication level, the literature appears large and highly consistent.
Data-provenance mapping You discover that nine of those 14 positive papers use the National Student Cohort during overlapping years. Three more papers use one university consortium database. The remaining eight papers across the entire review come from separately recruited populations.
Population assessment The nine National Student Cohort papers use different sample sizes and outcomes but contain substantial or probable participant overlap. They therefore do not represent nine independent replications.
Synthesis decision For one meta-analysis in which several cohort papers estimate essentially the same association, you apply your prespecified rule to select the estimate most appropriate to the target population and outcome. Other papers from the cohort remain available for different outcomes and descriptive analyses.
Sensitivity analysis You examine whether the overall interpretation changes when only one estimate from each substantially overlapping dataset family contributes to the relevant synthesis.

Your review can still report that many analyses have been conducted using the National Student Cohort. What changes is the interpretation: repeated analysis within that cohort is no longer presented as equivalent to repeated independent confirmation across nine populations.

05 · What Researchers Often Get Wrong

Common Mistakes When a Dataset Generates Much of the Literature

Misconception

Twenty Papers Mean Twenty Independent Replications

Not necessarily. Several publications may repeatedly analyze one cohort or database. Publication count measures research output, not the number of independent participant populations.

Misconception

Every Paper From the Same Dataset Should Be Collapsed Into One

This goes too far in the opposite direction. Papers may use non-overlapping samples, different outcomes, or distinct waves and may legitimately contribute to different parts of a review. Assess the actual population relationship and synthesis question.

Misconception

The Largest Analysis Should Always Be Retained

Sample size is only one consideration. A smaller analysis may better match the target population, outcome, exposure definition, follow-up period, or methodological requirements of the review.

Misconception

Different Authors Guarantee Independent Evidence

Public cohorts and large databases may be analyzed by many unrelated research groups. Independence comes from the underlying observations, not the names on the publication.

Misconception

This Problem Matters Only When Participants Are Duplicated Exactly

Partial overlap also creates dependence. Samples can share some participants while differing in years, eligibility restrictions, outcomes, or subgroups.

Misconception

Dataset Dominance Is Only a Meta-Analysis Problem

No. A narrative review can also overstate consistency or replication if numerous publications from one dataset are described as though they were independent confirmations.

06 · What This Means for You

Make Data Provenance Part of the Evidence Structure

Add the underlying data source and analytic population to your extraction framework whenever repeated dataset use is plausible. Then examine dataset families before drawing conclusions from publication counts or combining estimates.

A simple decision framework

If several papers use the same dataset but clearly non-overlapping populations
They may contribute as independent samples when the design and analysis support that interpretation.
If several papers use substantially overlapping participants for the same synthesis
Do not automatically treat their effect estimates as independent; apply an appropriate selection or analytical strategy.
If papers from one dataset address different review outcomes
Allow them to contribute where relevant while preserving their common data provenance.
If participant overlap is probable but cannot be quantified
Document the uncertainty and consider sensitivity analyses that reduce repeated representation from the dataset.
If one dataset accounts for a large share of the literature
State this explicitly and distinguish consistency within that resource from replication across independent populations.

The central discipline is interpretive as much as statistical. Ask not only “How many papers support this finding?” but also “How many substantively independent populations have tested it?” Those questions can produce very different answers.

07 · A Quick Checklist

How to Keep a Prolific Dataset in Proportion

Before interpreting the size and consistency of the evidence base, check:
Which cohort, registry, survey, biobank, administrative database, health system, or other data resource underlies each paper.
How many included publications arise from each major dataset family.
Whether those papers use overlapping years, waves, locations, eligibility criteria, or participant groups.
Whether samples are completely overlapping, nested, partially overlapping, plausibly independent, or uncertain.
Whether several estimates from the same data resource are entering the same synthesis under an independence assumption.
Whether any rule for choosing among overlapping estimates is based on prespecified review-relevant criteria rather than whichever result is strongest.
Whether sensitivity analyses are warranted to test the influence of heavily reused datasets.
Whether the narrative distinguishes repeated analyses within datasets from replication across independent populations.
Whether readers can see clearly when a substantial proportion of the literature derives from one or a few data resources.
08 · Frequently Asked Questions

Questions About Dataset Dominance in Evidence Synthesis

Should I include only one paper from each dataset?

No universal rule requires this. Papers from the same dataset may use independent samples, address different outcomes, or provide information relevant to different syntheses. The decision should depend on participant overlap and the inference being made.

Can several papers from one dataset all be valid studies?

Yes. Dataset reuse is not inherently invalid. The concern is whether their shared data provenance is recognized when interpreting replication or combining estimates that may not be independent.

How can I tell whether one dataset dominates my review?

Group included papers by their underlying data source and compare the number of publications, analytic populations, and outcomes contributed by each resource. A simple dataset-provenance table often makes concentration immediately visible.

Should I choose the largest paper when several samples overlap?

Not automatically. Size may be relevant, but the selected estimate should also fit the review's population, outcome, exposure or intervention, follow-up, and methodological criteria. Any selection rule should be transparent and preferably prespecified.

What if I cannot determine exactly how much the samples overlap?

Record the uncertainty rather than inventing an overlap percentage. Dataset, observation period, sites, and eligibility criteria may still establish plausible dependence, and sensitivity analyses can assess whether conclusions depend on including all potentially overlapping estimates.

Can one dataset contribute to several different meta-analyses?

Yes. Different papers may provide different eligible outcomes or time points. The key issue is whether overlapping participant contributions are being treated as independent within the particular synthesis, not whether the dataset appears elsewhere in the review.

Does repeated analysis of one dataset provide replication?

It can provide evidence of analytical robustness or reproducibility within that data resource, depending on how independently the analyses were conducted. It should not automatically be described as equivalent to replication across independently recruited populations.

How should I describe a literature dominated by one cohort?

State the concentration directly. Report how many analyses arise from the cohort, whether their participant samples overlap, and how findings compare with evidence from independent populations. This is more informative than relying on the total number of publications.

09 · The Bottom Line

Count Where the Evidence Comes From, Not Only How Many Papers Exist

The Bottom Line

Prevent a prolific dataset from dominating your interpretation by mapping the data source behind every relevant paper, identifying overlapping analytic populations, and distinguishing repeated analyses within a dataset from replication across independent populations.

Do not automatically exclude every additional paper from a shared dataset. Instead, match the response to the actual dependence: retain distinct information where it is useful, avoid treating overlapping estimates as independent, test influential decisions when appropriate, and report clearly when much of the apparent literature originates from the same underlying population.

10 · Sources and Further Reading

Authoritative and Methodological Sources on Dataset Reuse and Double-Counting

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes