01 · The Question
Can Statistics Recover Studies You Cannot Observe?
You suspect that a meta-analysis overestimates an effect because unfavorable studies were less likely to be published. You cannot identify those studies, do not know how many exist, and have no access to their results.
Could a statistical method nevertheless correct the meta-analysis and tell you what the result would have been if every study had been available?
Methods exist for exploring the consequences of missing evidence, and some attempt to estimate effects after accounting for possible publication processes. But there is a fundamental limitation: the missing studies are unknown. Their number, size, methods, precision, and results are not observed.
Statistics can model that uncertainty. They cannot make the missing research cease to be unknown.
03 · What You Need to Know
Why Missing Research Cannot Simply Be Reconstructed
The fundamental problem is that the missing evidence is unobserved
Suppose a meta-analysis contains 20 published studies. You suspect that additional studies exist because statistically significant findings may have been more likely to reach publication.
You do not necessarily know whether two studies are missing or twenty. You may not know whether they were small pilot studies or large multicenter investigations. You do not know their effect estimates, standard errors, methodological quality, populations, or even whether they measured exactly the outcome being synthesized.
Many different hidden evidence bases can therefore be consistent with exactly the same observed literature.
Observed evidence
Studies and results you can identify, inspect, extract, and evaluate directly.
Unknown missing evidence
Potential studies or results whose number, characteristics, and findings are partly or completely unobserved.
A statistical procedure cannot uniquely determine which unseen evidence base actually existed without additional information or assumptions.
Known missing results are easier to investigate than unknown missing studies
There is an important difference between knowing that a study exists but lacking one of its results and suspecting that entire unidentified studies may be absent.
Cochrane notes that the impact of selective non-reporting within identified studies can generally be quantified more readily because the number of identified studies with missing results is known. By contrast, selective non-publication may involve an unknown number of studies.
For example, if a trial registry identifies six eligible trials and two have no accessible results, you at least know the missing studies exist. You may be able to contact investigators, retrieve regulatory documents, inspect registry results, or explore assumptions about those two specific studies.
If no comprehensive registry exists, there may be additional studies you do not even know to search for. That is a much less constrained problem.
Correction requires a model of why studies went missing
Methods intended to adjust for publication bias must make assumptions about the process determining whether research becomes observable.
For example, selection models can assume that publication probability depends on features such as a study's standard error, effect direction, magnitude, or P value. The observed studies are then analyzed under a model representing that selection process.
Cochrane describes selection models as methods developed to estimate intervention effects corrected for bias due to missing results, but emphasizes that some model parameters cannot be estimated precisely. Sensitivity analyses are therefore used to examine estimates across different assumptions about the severity of selection bias.
The resulting estimate is not a recovered historical record. It is an estimate conditional on the assumed selection mechanism.
Different assumptions can produce different corrected estimates
This is the central limitation of adjustment methods. If you assume that statistically nonsignificant small studies were much less likely to be published, you may obtain one adjusted estimate. A weaker selection mechanism can produce another. A different model of how publication decisions operate can produce another still.
If conclusions remain similar across a credible range of assumptions, that stability is reassuring. If estimates change dramatically when reasonable assumptions change, the observed conclusion is more vulnerable to missing evidence.
| Question |
What can sometimes be estimated? |
What remains uncertain? |
| How might selective publication affect the pooled estimate? |
Alternative estimates under specified selection models |
Whether the assumed publication mechanism is correct |
| Would plausible missing evidence change the conclusion? |
Sensitivity of the conclusion to specified scenarios |
Which scenario describes the actual missing studies |
| How many studies were never published? |
Some methods may infer patterns compatible with missing evidence |
The true number of unidentified studies generally remains unknown |
| What did the missing studies find? |
Possible distributions can be modeled |
Their actual results remain unknown unless recovered |
Funnel plots diagnose patterns, not invisible studies
A funnel plot can reveal relationships between study estimates and their precision that may be compatible with missing evidence. For example, smaller studies with unfavorable findings may appear to be absent from one region of the plot.
But asymmetry can arise for reasons other than publication bias, including genuine differences between smaller and larger studies, methodological differences, or other forms of heterogeneity. Cochrane explicitly cautions that funnel-plot asymmetry should not be treated as diagnostic of non-reporting bias.
The reverse is also important. A visually symmetrical funnel plot does not establish that every relevant study was published.
A funnel plot therefore provides evidence about the observed distribution of studies. It does not show you the actual studies that are absent.
Tests of funnel-plot asymmetry have the same fundamental limitation
Statistical tests such as Egger's test examine whether study effect estimates are associated with measures of study size or precision more strongly than expected by chance.
Such tests can identify small-study effects under appropriate conditions. They cannot uniquely determine why those effects exist. Publication bias is one possible explanation among several.
Cochrane also notes that tests for funnel-plot asymmetry are applicable only to a minority of meta-analyses and that their interpretation requires attention to alternative explanations.
Watch Out
A statistical test for publication bias is not a detector that observes unpublished studies. It evaluates patterns in the studies you can see and asks whether those patterns are compatible with particular forms of missingness.
Selection models are useful precisely because they expose the assumptions
Selection models formalize a hypothetical relationship between study results and the probability that those results appear in the evidence base.
One advantage is that they allow researchers to ask a more informative question than "Is there publication bias?" You can instead ask how the pooled estimate behaves as the assumed degree of selective publication becomes stronger.
Cochrane notes that if estimates remain relatively stable across assumed selection models, the unadjusted estimate may be comparatively robust to non-reporting bias under those assumptions. If estimates vary substantially, missing evidence may plausibly drive the observed result.
However, selection models themselves can be misleading if their assumptions are wrong. In particular, apparent small-study effects may arise through mechanisms other than non-reporting bias.
Regression-based adjustments also depend on assumptions
Other approaches place greater emphasis on larger studies or extrapolate relationships between study size and estimated effects. These methods can provide useful alternative estimates when their assumptions are reasonable.
They still do not reveal the contents of missing studies. Cochrane notes that regression-based approaches generally require enough studies to estimate the relationship appropriately and can behave poorly when all available studies are small.
The existence of an adjusted number should therefore not be confused with certainty that the adjustment recovered the truth.
The best solution is often to recover evidence rather than statistically invent it
If missing results can actually be obtained, they provide more information than a model guessing what those results might have been.
Cochrane recommends searching beyond journal publications when relevant, including trial registries, regulatory sources, manufacturers, sponsors, and study authors. Its guidance notes examples in which inclusion of previously unavailable evidence materially changed conclusions, although the impact is not always dramatic.
This creates a useful hierarchy. First, try to identify and retrieve the missing evidence. If important uncertainty remains, assess risk of bias and examine robustness under plausible assumptions.
Prospective systems can prevent a problem that retrospective statistics cannot fully solve
The deepest solution to unknown unpublished research is not a more ingenious correction after the evidence disappears. It is creating systems that make the existence and results of studies observable in the first place.
Prospective registration, protocols, results reporting, repositories, regulatory disclosure, and prospective meta-analysis can reduce uncertainty about which studies were initiated and what they found.
Cochrane describes inception cohorts as one strategy in which studies are identified from a set known to have been initiated, such as prospectively registered trials. This allows reviewers to account for studies with and without available results rather than guessing how many unidentified studies might exist.
Correction and sensitivity analysis should not be treated as the same claim
The language researchers use matters here.
"We corrected publication bias"
This can imply that the missing evidence has been accurately reconstructed and the bias removed.
"We examined sensitivity to plausible publication-bias mechanisms"
This states more accurately that conclusions were evaluated under specified assumptions about missing evidence.
The second statement usually better reflects what these methods can establish when the underlying unpublished evidence remains unknown.
This connects directly to the question of whether missing studies could plausibly change the conclusion. You often do not need to know the exact hidden evidence to learn something useful. Knowing that the conclusion collapses under modest plausible missingness, or remains stable under severe credible scenarios, can itself inform interpretation.