01 · The Question
Are Separate Papers Necessarily Independent Studies?
You find four papers investigating the same relationship. They have different titles, were published in different journals, and perhaps even have somewhat different author lists. It is easy to read them as four independent pieces of evidence.
Then you notice something familiar: the same university, similar recruitment dates, nearly identical sample characteristics, or the name of the same longitudinal project buried in the Methods section.
The four papers may indeed report different analyses, but they may not represent four independent collections of evidence. Distinguishing publications from underlying studies is essential when deciding how much replication actually exists.
03 · What You Need to Know
How Apparently Separate Studies Can Be Connected
A Paper Is Not the Same Thing as a Study
This distinction is fundamental in evidence synthesis.
One research project can generate several journal articles. A clinical trial might produce a primary outcomes paper, a secondary outcomes paper, a long-term follow-up, and subgroup analyses. A longitudinal cohort might support dozens or even hundreds of publications addressing different questions.
There is nothing inherently problematic about this. Researchers should often extract as much legitimate knowledge as possible from valuable data.
The problem begins when readers mistake several reports arising from the same underlying study for several independent studies. Cochrane explicitly treats studies rather than reports as the unit of interest in systematic reviews and requires multiple reports of the same study to be linked together. It warns that duplicate inclusion can bias a synthesis.
Publication or report
A document communicating some aspect of a study, such as a journal article, conference abstract, follow-up report, or secondary analysis.
Study
The underlying investigation from which one or more reports may arise.
The Same Dataset Can Generate Apparently Different Studies
Overlap is not always obvious. Two papers may ask different research questions, analyze different variables, or use different subsets of participants while drawing observations from the same larger dataset.
For example, one paper might examine academic performance, another mental well-being, and another technology use among participants in the same longitudinal cohort. These are legitimately different analyses. Yet if all three provide estimates relevant to your synthesis, the estimates may be statistically dependent because some or all of the same participants contributed data to them.
Cochrane specifically notes that effects calculated from multiple outcomes or measures involving the same participants are statistically dependent.
Partial Overlap Matters Too
Dependence is not always all or nothing.
Suppose Study A analyzes 2,000 participants from a national survey. Study B analyzes 1,200 participants from the same survey but applies different eligibility criteria. Study C combines one wave of that survey with newly collected data.
None is an exact duplicate of another. Yet neither are they fully independent.
The degree of overlap matters because repeatedly counting some participants can give their observations disproportionate influence. In quantitative synthesis, failure to account for this dependence can also make estimates appear more precise than the underlying information warrants. Cochrane warns, for example, that double-counting participants can spuriously increase precision.
Shared Research Groups Are a Clue, Not Proof of Dependence
Seeing the same authors across several papers should prompt investigation, but it does not prove that the studies share data.
A productive research group might conduct five genuinely independent experiments. Conversely, two papers with substantially different author lists might analyze the same large public dataset.
Author overlap is therefore diagnostic information rather than a verdict. Cochrane lists common authors among several clues that can help identify multiple reports from the same study, alongside study location, intervention details, participant numbers, baseline characteristics, and study dates.
Public Datasets Can Create Less Obvious Dependence
Shared data are particularly easy to overlook when researchers analyze large public, administrative, or longitudinal datasets.
Two teams that have never collaborated may independently publish analyses based on the same national survey, health registry, educational database, biobank, or cohort. Their intellectual and analytical decisions may be independent, while their observations are not.
This distinction is useful because independence has several dimensions. Different researchers may provide analytical independence without providing participant-level independence.
| Type of independence |
Question to ask |
Why it matters |
| Participant independence |
Are the observations from different people or units? |
Repeated participants can create statistical dependence. |
| Dataset independence |
Were the data collected separately? |
Different analyses of one dataset are not equivalent to new data collection. |
| Research-team independence |
Were different investigators responsible for the studies? |
Independent teams may reduce dependence on one group's decisions or practices. |
| Methodological independence |
Do the studies rely on different measurements, designs, or assumptions? |
Different methods can challenge explanations tied to one methodological weakness. |
Different Research Questions Do Not Make the Observations Independent
A paper can be genuinely novel without supplying a new independent sample.
Suppose researchers collect a rich dataset containing demographics, educational outcomes, psychological measures, and technology-use variables. One paper studies achievement, another engagement, and another well-being. Each may make an original contribution.
But if you later synthesize several estimates from those papers to answer one broader question, their shared participants still matter. Novelty of the research question and independence of the evidence are separate issues.
Why Double-Counting Can Distort a Synthesis
Imagine a particularly large study that generates four publications. If a reviewer accidentally treats those publications as four independent studies, the underlying dataset effectively enters the evidence base several times.
If all four analyses point in the same direction, the literature can appear more consistently supportive than it really is. In a meta-analysis, inappropriate reuse of participant data can also distort weighting and precision.
Cochrane therefore requires multiple reports from one study to be collated and warns that it is wrong to treat those reports as multiple studies.
Overlap Can Also Occur Between Systematic Reviews
The same problem appears one level higher. Two systematic reviews may look like independent summaries while containing many of the same primary studies.
If an overview combines those reviews without accounting for overlap, some primary studies can effectively influence the conclusion repeatedly. Cochrane identifies this as a specific problem in overviews of reviews because double-counting the same primary-study outcome data gives those studies excessive influence.
How Can You Detect Shared Data?
Do not rely only on titles and abstracts. Inspect the Methods sections and supplementary information.
Look for named cohorts, trial registration numbers, dataset names, recruitment locations, institutions, recruitment dates, sample sizes, baseline characteristics, intervention details, funding projects, and overlapping author lists. Cochrane recommends several of these clues when identifying multiple reports and notes that resolving uncertain cases can require contacting investigators.
Sometimes the relationship is explicit: “This study is a secondary analysis of...” Other times, some detective work is unavoidable. Evidence synthesis occasionally resembles genealogy with confidence intervals.
Shared Data Do Not Make the Papers Useless
Once you discover overlap, do not automatically discard every secondary paper.
Secondary reports may contain additional outcomes, methodological details, follow-up information, or analyses absent from the primary report. Cochrane specifically advises retaining relevant secondary reports and collating information across them rather than simply throwing duplicates away.
The objective is to represent the underlying evidence correctly.
This is also why convergence of evidence should be distinguished from repetition of the same evidence. Multiple useful analyses can enrich understanding without constituting multiple independent replications.
07 · A Quick Checklist
How to Check Whether Studies Are Actually Independent
Before counting papers as independent evidence, check:
Compare the names of cohorts, surveys, trials, databases, registries, and research projects.
Compare recruitment locations, institutions, and data-collection dates.
Compare sample sizes and baseline characteristics for suspiciously similar participant profiles.
Look for statements identifying a paper as a secondary analysis, follow-up, substudy, or analysis of previously collected data.
Check trial registration numbers, dataset identifiers, grant numbers, and project names where available.
Use author overlap as a clue, but do not assume that the same authors mean the same data or that different authors mean different data.
Determine whether samples overlap completely, partially, or not at all.
In quantitative synthesis, use an analysis that appropriately handles dependent estimates rather than double-counting participants.