03 · What You Need to Know
Where does uncertainty in research evidence come from?
Uncertainty belongs to a specific conclusion
It is rarely useful to say simply that “the evidence is uncertain.” Uncertain about what?
You may be uncertain whether an intervention produces any meaningful benefit. Or you may have strong evidence that it produces some benefit but remain uncertain whether the effect is small or large. You may know the short-term effect reasonably well while having little evidence about durability. You may have credible evidence in adults but uncertain applicability to adolescents.
Broad uncertainty
“We do not know whether this intervention works.”
Specified uncertainty
“The available evidence suggests a short-term benefit, but its magnitude and persistence beyond six months remain uncertain.”
The second statement tells the reader what the evidence does and does not establish.
Risk of bias creates uncertainty about whether the estimate is trustworthy
Bias concerns systematic distortion.
If participants were selected in ways related to the outcome, important confounders were not adequately addressed, outcome measurement differed systematically between groups, missing data were substantial, or results were selectively reported, the observed estimate may differ from the quantity the study intended to estimate.
In the GRADE approach used by Cochrane, risk of bias is one of five principal domains considered when assessing certainty in a body of evidence.
This form of uncertainty asks: Would the result look meaningfully different if the relevant biases were absent?
Imprecision means the data permit materially different conclusions
Imprecision is often easier to see because it appears in wide confidence intervals or similarly broad uncertainty intervals.
Suppose an intervention's estimated effect is beneficial, but the confidence interval is compatible with almost no meaningful benefit and with a substantial benefit. The central estimate alone does not resolve which interpretation is closer to reality.
Cochrane's GRADE guidance recommends judging imprecision by whether confidence intervals encompass substantively different possibilities, such as little or no effect and important benefit or harm, rather than merely whether a conventional significance threshold is crossed.
Watch Out
A non-significant result is not synonymous with “no effect,” and a significant result is not synonymous with “certainty.” Ask which effect sizes remain reasonably compatible with the evidence and whether those possibilities would lead to different substantive conclusions.
Indirectness creates uncertainty about whether evidence transfers to your question
You can be highly confident about what a study found while remaining uncertain about whether it answers your question.
Perhaps the available trials involve adults while your population is adolescents. Perhaps an intervention was tested under intensive researcher supervision but will be implemented routinely without that support. Perhaps studies measured a surrogate outcome rather than the outcome you actually care about.
GRADE explicitly evaluates indirectness by considering whether the available population, intervention, comparator, and outcome sufficiently match the question being asked.
Indirectness therefore separates two judgments: confidence in the observed evidence and confidence in applying that evidence elsewhere.
Unexplained inconsistency creates uncertainty about what effect to expect
If credible studies produce materially different estimates and the differences remain unexplained, a single general conclusion becomes harder to defend.
Cochrane's GRADE guidance states that when important heterogeneity affects interpretation and no plausible explanation can be identified from the available data, certainty should be reduced.
But inconsistency should not be declared merely because numbers differ. First investigate why important studies disagree and whether different questions, populations, measures, or methods explain the apparent conflict.
Sparse evidence can create uncertainty, but study count is not the real issue
Researchers often say “only three studies exist, therefore the evidence is weak.” The number of studies by itself is not the key quantity.
One very large, rigorous study can sometimes provide more precise evidence than ten tiny studies. Conversely, many studies can still leave considerable uncertainty if they are biased, indirect, or based on very few relevant events.
Cochrane's GRADE guidance explicitly advises against using the number of studies itself as the reason for judging imprecision. Instead, the amount of information and the width and substantive implications of confidence intervals should be considered.
Publication bias creates uncertainty about the evidence you cannot see
Your synthesis can only analyze evidence you have identified. If the probability of publication or reporting depends on the results, the visible literature may provide an incomplete and systematically distorted picture.
Publication bias is therefore one of the five domains considered by GRADE when assessing certainty.
This is one reason grey and unpublished evidence may matter for some questions. The uncertainty does not arise merely because something might theoretically be missing. It arises when there are credible reasons to suspect that missing evidence could materially change the conclusion.
Dependence on one study or dataset creates structural uncertainty
A conclusion can look well established because many papers discuss it while ultimately depending on one underlying study, cohort, research group, or dataset.
That creates a different form of vulnerability. You may understand that particular dataset extremely well while knowing little about whether the result survives new samples, settings, investigators, or methods.
Ask whether important conclusions depend heavily on one evidential source. Lack of independent corroboration does not automatically invalidate a finding, but it should limit claims about robustness and generalizability.
Different uncertainties can coexist
Evidence is rarely uncertain for only one reason.
| Source of uncertainty |
Question it raises |
| Risk of bias |
Would a less biased study produce a meaningfully different result? |
| Imprecision |
Which substantively different effect sizes remain compatible with the evidence? |
| Indirectness |
Does the available evidence apply closely enough to the actual question? |
| Inconsistency |
Why do credible studies differ, and what effect should be expected across settings? |
| Publication bias |
Could missing evidence materially change the visible pattern? |
| Limited independence |
Would the finding survive a genuinely new sample, group, dataset, or method? |
| Missing dimensions |
Are important outcomes, populations, time periods, harms, or mechanisms simply unstudied? |
A strong synthesis identifies which of these matters for each major conclusion rather than compressing them into one generic sentence about limitations.
Uncertainty about magnitude is different from uncertainty about direction
Suppose ten credible studies all indicate a beneficial effect, but estimates vary from very small to moderate. You may have relatively little uncertainty about direction but substantial uncertainty about magnitude.
Conversely, estimates may range from meaningful benefit to meaningful harm. That is a much more consequential form of uncertainty because different conclusions remain plausible.
These distinctions matter when translating evidence into decisions. “Probably beneficial, magnitude uncertain” communicates something very different from “benefit or harm both remain plausible.”
Uncertainty about generalizability is different from uncertainty about the observed effect
A highly controlled study may provide precise evidence about what happened under its study conditions. The uncertainty may lie in whether the result applies elsewhere.
For example, an educational intervention tested by its developers in well-resourced universities with intensive implementation support may work convincingly there. Whether the same effect occurs in institutions with larger classes, less training, or different student populations is a separate question.
Do not reduce confidence in the original finding simply because generalizability is uncertain. Locate the uncertainty correctly.
Uncertainty is not evidence of absence
If the evidence cannot establish whether an effect exists, you do not automatically have evidence that the effect is absent.
This becomes especially important with imprecise estimates. A study whose interval includes meaningful benefit, no meaningful effect, and meaningful harm has not demonstrated equivalence among those possibilities. It has failed to distinguish them adequately.
The distinction between absence of evidence and evidence of absence prevents uncertainty from being converted into an unsupported negative conclusion.
Uncertainty is also not a license to say anything is possible
The opposite mistake is treating imperfect evidence as though every conceivable explanation remains equally plausible.
Evidence can narrow uncertainty substantially without eliminating it. Perhaps a very large harmful effect is implausible while small benefit, no meaningful effect, and small harm remain possible. Perhaps the evidence strongly supports an association but leaves causality unresolved.
Scientific caution should be calibrated to what remains genuinely open, not expanded until every sentence dissolves into “more research is needed.”
Certainty should be assessed at the level of outcomes or conclusions
GRADE assesses certainty for a body of evidence concerning a particular outcome, not by assigning one global quality label to an entire literature. Cochrane describes certainty as the extent to which one can be confident that an estimate of effect or association is close to the quantity of specific interest.
This outcome-specific approach is useful conceptually even when you are not conducting a formal GRADE assessment. A field can have strong evidence about one outcome and highly uncertain evidence about another.