01 · The Question
How many genuinely independent foundations are holding up the conclusion?
A research claim appears everywhere. Twenty papers discuss it. Several reviews summarize it. Citation counts are substantial. At first glance, the evidence base looks broad.
Then you map where the evidence comes from.
Eight papers use the same longitudinal dataset. Five come from one research group. Several use the same measurement instrument. Most studies recruit from similar institutions. Nearly every analysis uses the same observational design.
The literature may still contain valuable evidence. But its apparent breadth exceeds its evidential diversity.
The question is therefore not merely how many publications support the conclusion. It is how much of that conclusion would remain if its dominant study, dataset, research group, population, instrument, or methodological approach disappeared.
03 · What You Need to Know
How can a large literature rest on a surprisingly narrow foundation?
Publication count is not evidence count
A single research project can generate numerous legitimate publications. One paper may report primary outcomes, another secondary outcomes, another a subgroup analysis, another long-term follow-up, and several others distinct research questions using the same dataset.
All may contribute useful information. They do not automatically provide independent replication.
Cochrane explicitly states that studies rather than reports are the principal unit of interest in systematic reviews and requires multiple reports from the same study to be collated. It also warns that treating multiple reports as multiple studies is wrong because the same underlying investigation can otherwise be counted repeatedly.
This is why the first step is to distinguish publications from independent studies.
Dataset dependence can hide behind different research questions
Dependence is not limited to obvious duplicate publication.
A large public or proprietary dataset may support dozens or hundreds of papers asking different questions. Those papers are not duplicates. But when several are cited as corroboration for the same broad conclusion, their shared data source matters.
If the dataset has a particular sampling limitation, measurement problem, missing-data pattern, or contextual peculiarity, multiple analyses may inherit that vulnerability.
Analytical replication
Different analyses, models, or questions are examined using the same or substantially overlapping underlying data.
Independent empirical replication
A finding is tested using genuinely new observations, participants, settings, or datasets capable of providing independent corroboration.
Both can be useful. They answer different robustness questions.
Research-group dependence can preserve shared assumptions
A productive research group may legitimately dominate an emerging field. Its members may conduct several independent studies rather than repeatedly analyze one dataset.
Even then, independence is not complete in every relevant sense.
The same group may use similar recruitment strategies, intervention procedures, instruments, theoretical assumptions, coding practices, analytical pipelines, or laboratory conditions. A finding repeatedly obtained by one team is informative, but replication by independent investigators tests whether the result survives a different set of implementation decisions and researcher-specific practices.
Research-group concentration should therefore be described rather than treated as an accusation. Expertise can produce excellent programs of research. The evidential question is whether the conclusion has also survived genuinely independent attempts to test it.
Method dependence can make a phenomenon look more universal than it is
Suppose nearly every study supporting a claim uses cross-sectional self-report surveys. The studies may involve different participants and institutions, yet the evidence still depends heavily on one methodological family.
If the association partly reflects common-method bias, recall error, social desirability, reverse causation, or another vulnerability characteristic of that approach, independent samples alone may not resolve the problem.
Evidence becomes more robust when a conclusion survives defensible methods with meaningfully different assumptions and biases, provided those methods actually address comparable questions.
| Type of dependence |
Question to ask |
| Study dependence |
How many publications arise from the same underlying investigation? |
| Dataset dependence |
How many findings rely on the same participants, cohort, archive, registry, or database? |
| Research-group dependence |
How much evidence comes from investigators sharing similar procedures, assumptions, or implementation practices? |
| Population dependence |
Would the conclusion survive beyond the narrow populations repeatedly studied? |
| Measurement dependence |
Does the conclusion rely heavily on one instrument or operational definition? |
| Method dependence |
Has the claim survived methods with different strengths and vulnerabilities? |
| Analytical dependence |
Do findings depend on one modeling strategy, specification, threshold, or analytical convention? |
Population dependence limits generalization even when replication is genuine
Imagine twelve independent studies from twelve universities, all conducted among undergraduate psychology students in highly resourced institutions.
The studies are genuinely independent at the sample level. Yet the broader conclusion may still depend heavily on one type of population and setting.
This is a form of indirectness when the intended claim concerns a broader population. GRADE guidance recognizes that evidence may be indirect when studies address a restricted version of the population, intervention, comparator, or outcome relevant to the broader question.
The evidence might strongly establish what happens in the studied population while providing weaker grounds for universal generalization.
Measurement dependence can make a construct look better established than it is
Suppose twenty independent studies report a relationship between two constructs, but all measure one construct using the same questionnaire.
The evidence may robustly establish an association involving scores on that instrument. Whether it establishes the broader theoretical construct depends partly on the validity of the measure.
If the instrument systematically captures something narrower or different from the intended construct, independent samples will repeatedly reproduce the same measurement assumption.
Ask whether the finding survives alternative defensible operationalizations.
Analytical dependence can be less visible than dataset dependence
Studies can use independent datasets yet all apply essentially the same analytical model and assumptions.
Perhaps a conclusion appears only after dichotomizing a continuous variable at one conventional threshold. Perhaps every study adjusts for the same questionable set of covariates. Perhaps one preprocessing pipeline dominates an imaging literature.
When plausible alternative analyses produce materially different conclusions, analytical dependence becomes part of the evidential uncertainty.
One highly influential study can shape an entire research program
A foundational study can determine the theory, measurement, design, and hypotheses adopted by later work. Subsequent papers may look independent because they collect new samples, yet still reproduce the conceptual architecture of the original study.
This is not necessarily a flaw. Scientific programs develop cumulatively.
But if the foundational assumptions are wrong, downstream research can inherit the problem. This is one reason consequential claims should be traced back to their original evidence rather than judged only by how frequently later papers repeat them.
Dependence changes what “consistent evidence” means
Ten studies from one dataset can be internally consistent. Five independent datasets analyzed by one group can also be consistent. Five independent teams using different credible methods can provide another form of consistency.
Those patterns should not be described as though they provide identical corroboration.
When evaluating whether evidence is genuinely consistent, ask how many opportunities the conclusion has had to fail under genuinely new data, settings, investigators, measurements, and methods.
Dependence is not automatically weakness
A uniquely valuable dataset may be the only ethical or feasible way to study a rare population. One research group may possess specialized expertise unavailable elsewhere. A validated instrument may appropriately dominate measurement because it performs better than alternatives.
The issue is not that concentration exists. The issue is whether your conclusion acknowledges what has and has not been independently tested.
Watch Out
Do not punish a literature merely for having a successful research program or an excellent dataset. Evidence concentration becomes problematic when publication volume is mistaken for independent corroboration or when claims extend beyond what the concentrated evidence has actually tested.
A simple removal test can reveal dependence
One practical diagnostic is to imagine removing the dominant source.
If the largest study disappeared, what would remain? If every paper using one dataset were removed, would the conclusion still stand? If one research group's work vanished, would independent investigators still support the claim? If one measurement instrument were excluded, would alternative operationalizations converge?
This is not necessarily a formal sensitivity analysis, although formal syntheses can perform analogous analyses quantitatively. It is a conceptual stress test for the structure of the literature.
Dependence affects certainty more than truth
A conclusion heavily dependent on one excellent study may ultimately be correct. The problem is that fewer independent tests have challenged it.
Therefore, dependence should usually modify your confidence and wording rather than automatically reverse the conclusion.
Instead of writing “numerous studies establish X,” you may write that evidence repeatedly supports X but remains concentrated in one dataset or research program and awaits broader independent replication.
That is not rhetorical weakness. It is a more accurate description of the evidence.