03 · What You Need to Know
Why a Highly Productive Dataset Can Distort the Apparent Evidence Base
Start by Separating Papers, Analyses, Datasets, and Populations
Four quantities can easily become conflated: the number of publications, the number of analyses, the number of underlying datasets, and the number of independent participant populations. They are not interchangeable.
A national cohort might support 20 eligible papers. Those papers could contain 20 distinct analyses while drawing repeatedly from one participant pool. Another literature containing 20 papers from 20 independently recruited populations has a very different evidential structure even though the publication count is identical.
Publication volume
How many papers or reports exist.
Independent population evidence
How many substantively independent participant populations contribute information relevant to the question.
This distinction follows the broader principle that systematic reviews should not equate reports with studies. Cochrane emphasizes that studies rather than individual reports are the primary units of interest and that multiple reports of one study should be identified and linked. The same logic becomes more complicated when nominally separate studies reuse a common database or cohort.
A Prolific Dataset Is Not Inherently a Problem
Large cohorts, biobanks, administrative databases, registries, longitudinal surveys, and electronic health-record systems exist partly so that they can support multiple investigations. Reanalysis can answer new questions, examine different outcomes, test subgroups, extend follow-up, and improve understanding of a population.
Restricting every dataset to one publication would therefore be methodologically crude. Two papers from the same resource may analyze mutually exclusive participants, different waves, or genuinely different outcomes relevant to separate syntheses.
The concern arises when repeated use of the resource is mistaken for independent replication or when overlapping estimates enter an analysis under an assumption of independence.
First Establish How Much Dataset Reuse Exists
Data provenance should be extracted explicitly. For every included paper, record the underlying cohort, registry, survey, database, biobank, health system, or other data source whenever it can be identified.
Then record the features needed to reconstruct each analytic population:
- recruitment or observation years;
- study waves or releases;
- geographic and institutional coverage;
- age or population restrictions;
- eligibility and exclusion criteria;
- exposure or intervention definitions;
- outcome requirements;
- analytic sample size; and
- subgroup restrictions.
This converts an apparently flat list of papers into a map of where the evidence actually came from.
Same Dataset Does Not Necessarily Mean Same Participants
A shared data source is a warning to investigate overlap, not proof that two papers contain identical people.
A national database might contain mutually exclusive geographic regions. A longitudinal cohort may recruit new waves of participants. Two papers could use non-overlapping age groups or calendar periods. In such situations, the analyses may be substantially or completely independent despite sharing the same infrastructure.
Conversely, slightly different eligibility criteria can produce two samples that overlap extensively without being identical. This is why you need to determine whether multiple papers from the same dataset actually represent independent evidence.
Map Dataset Families Before Looking at the Overall Pattern
A useful organizational strategy is to group papers by their underlying data source before interpreting how many independent sources support a conclusion.
| Dataset Family |
Publications |
Population Relationship |
Interpretive Implication |
| National Cohort A |
10 papers |
Substantial overlap across waves and subgroups |
Many analyses, but not 10 independent replications |
| Hospital Database B |
4 papers |
Overlapping calendar periods |
Potential dependence requires further assessment |
| Survey C |
3 papers |
Mutually exclusive survey waves |
May provide more independent samples if participants do not recur |
| Independent studies |
8 papers |
Separately recruited populations |
Provide distinct population-level evidence |
The point is not to assign a simplistic weight to each row. It is to prevent ten papers generated from one cohort from visually masquerading as ten unrelated populations.
Do Not Count Repeated Analyses as Independent Replication
Replication usually carries a stronger implication than repeated analysis. If several research teams analyze overlapping samples from the same cohort and obtain similar findings, that may demonstrate robustness to certain analytical choices or research teams. It does not provide the same evidence about transportability across populations as reproducing the finding in independently recruited samples.
Within-dataset corroboration
Related analyses from the same data resource produce compatible findings under different questions, models, subgroups, or research teams.
Across-population replication
Compatible findings arise in participant populations whose observations are substantively independent of the original dataset.
Both forms of evidence can matter. They simply answer somewhat different questions about robustness.
Dataset Dominance Can Affect Narrative Reviews Even Without Meta-Analysis
Suppose 14 of 18 papers report a positive association. That initially sounds like broad consistency. If 11 of those positive papers use overlapping samples from one cohort while three independent studies produce mixed findings, the evidential picture is more complicated.
A narrative synthesis that simply counts positive and negative papers would give the prolific cohort considerable implicit weight. Even without a formal statistical model, publication frequency can shape the reader's impression of consistency.
Report the clustering directly. Instead of saying that “14 studies independently found the association,” describe how many analyses came from each major data source and how consistently the finding appeared across genuinely distinct populations.
Dataset Dominance Can Also Affect Meta-Analysis
The problem becomes statistical when estimates based on overlapping participants are entered into a meta-analysis as though they were independent.
Hussein and colleagues describe double-counting as the inclusion of the same individuals multiple times in one evidence synthesis and identify reuse of the same database, overlapping analytical periods, and common treatment groups as mechanisms through which it can occur.
The consequences depend on the extent of overlap and the analysis, but ignoring dependence can give repeated observations inappropriate influence and may make pooled estimates appear more precise than warranted.
Simply Choosing the Largest Paper Is Not a Universal Solution
One practical response to overlap is to select a single analysis from a cluster. For example, Hussein and colleagues describe a case in which several studies used UK Biobank data for a related question; the review retained one study after considering publication status and sample size.
That illustrates one possible strategy, not a universal rule. The largest paper may use a less appropriate outcome, a different population, or an inferior time point for your review. The newest paper is not automatically preferable either.
If selection is necessary, criteria should be tied to the review question and preferably specified in advance. Relevant considerations might include population fit, outcome definition, follow-up, methodological suitability, completeness, and risk of bias rather than sample size alone.
One Dataset May Legitimately Contribute to Several Different Syntheses
A cohort could provide an eligible mental-health outcome in one paper and an academic outcome in another. If your review analyzes those outcomes separately, both reports may appropriately contribute to their respective syntheses.
The key question is not whether a dataset has already appeared somewhere in the review. It is whether the same or overlapping participants are being given repeated independent representation within the particular inference you are making.
This is why blanket rules such as “one publication per dataset” can solve one problem by creating another.
Consider Sensitivity Analyses When Dataset Dominance Could Change the Conclusion
When several eligible estimates derive from potentially overlapping populations and no single analytical solution is clearly preferable, sensitivity analysis can be informative.
For example, you might compare the primary synthesis with an analysis retaining only one appropriately selected estimate from each substantially overlapping dataset family. If the conclusion changes materially, the evidence is sensitive to how repeated dataset contributions are handled.
Hussein and colleagues demonstrated empirically that removing overlapping population papers could change pooled estimates, uncertainty, and heterogeneity in a real-world case study. That does not imply the same magnitude or direction of change in every review, but it shows that overlap need not be inconsequential.
Dataset Diversity Is Not the Same as Population Diversity
Even several technically distinct databases can represent similar populations. Conversely, one large multinational resource can contain heterogeneous populations. Counting dataset names therefore remains an imperfect proxy for evidential diversity.
Consider the substantive characteristics of the populations as well: geography, healthcare system, socioeconomic context, age, ethnicity where relevant and appropriately reported, institutional setting, period of data collection, and other factors that affect the applicability of the evidence.
The goal is not to maximize the number of dataset names. It is to understand where the evidence comes from and how broadly the finding has actually been examined.
Make Dataset Provenance Visible in Your Results
If one or two data resources generate a substantial fraction of the included literature, readers should be able to see that. Depending on the review, this information may appear in study-characteristics tables, evidence maps, narrative synthesis, supplementary tables, or sensitivity analyses.
Transparency is especially important because publication titles and author lists may otherwise conceal the shared source. A reader seeing ten citations should not have to perform bibliographic archaeology to discover that eight arose from one cohort.
Watch Out
Do not “correct” dataset dominance by automatically deleting all but one paper from every shared data source. The relevant unit depends on the review question, outcome, participant overlap, and synthesis. First map the dependence; then decide what methodological response it requires.