01 · The Question
What Should You Do When the Evidence Does Not Agree?
One study reports a substantial benefit. Another finds little difference. A third suggests the effect occurs only for certain participants. You now have an awkward choice: write a neat overall conclusion or acknowledge that the literature tells a more complicated story.
The neat conclusion is often easier to write. It may also be less accurate.
Disagreement among studies can reveal differences in populations, interventions, measurements, research designs, implementation, or contexts. It can also arise from bias, imprecision, or chance. Sometimes the available evidence does not allow you to determine the reason at all.
Your job is not to make every study agree. It is to determine whether the disagreement matters and, when it does, represent it honestly in the synthesis.
03 · What You Need to Know
Disagreement Can Tell You Something the Average Cannot
First decide whether the studies genuinely disagree
Two studies do not necessarily conflict simply because one reports a statistically significant result and the other does not. Their estimated effects may be very similar, with the difference in significance arising from sample size or precision.
Likewise, different numerical estimates may still support the same substantive interpretation. Before describing evidence as contradictory, compare the findings themselves, their uncertainty, and what each result would imply for your research question.
This is why it helps first to determine whether you are seeing meaningful inconsistency rather than ordinary variation across studies.
Do not hide disagreement inside a majority count
Suppose seven studies report a positive effect and three do not. “Most studies found a positive effect” may be factually correct, but it can conceal information that changes the interpretation.
What if the three conflicting studies are the largest? What if they use a more appropriate outcome measure? What if all seven positive studies come from one setting while the conflicting studies come from another? What if the positive findings concern short-term outcomes and the others examine longer-term effects?
A study count does not answer these questions.
Watch Out
Do not treat the majority of studies as though research synthesis were a ballot. Studies differ in relevance, precision, design, execution, and evidential contribution. Numerical dominance does not automatically determine the most defensible conclusion.
Ask whether population differences could explain the pattern
An intervention may work differently for different populations. Age, baseline characteristics, prior knowledge, disease severity, socioeconomic circumstances, institutional characteristics, or other factors may modify an effect or relationship.
If studies involving one population repeatedly produce different findings from studies involving another, that pattern may deserve investigation. But do not assume that a visible subgroup difference proves effect modification. Formal subgroup analyses require appropriate comparisons, and post hoc explanations are particularly vulnerable to chance findings.
Population differences may also indicate that your overall conclusion has narrower limits of applicability than its wording initially suggests.
Check whether supposedly similar interventions or exposures are actually similar
Broad labels can conceal substantial variation. “Online learning,” “peer feedback,” “mindfulness,” “AI-assisted instruction,” or “exercise intervention” can describe quite different experiences.
Duration, intensity, implementation, content, instructor involvement, adherence, technological features, and co-interventions may differ enough to influence outcomes. If effects change alongside these characteristics, disagreement may reflect genuine differences in what participants received rather than failure to replicate the same intervention.
Outcome definitions and measurement can create apparent disagreement
Studies can examine the same broad construct while measuring it differently. One study may use a validated performance assessment, another a self-report scale, and another an administrative indicator. Follow-up periods may also differ.
If those measures capture different aspects of the phenomenon, their findings need not coincide. Before treating results as contradictory, ask whether the studies actually estimated the same outcome at a comparable time point.
Study design and methodological quality may help explain disagreement
Findings may differ because studies vary in susceptibility to bias. Randomized and non-randomized studies, prospective and retrospective designs, adjusted and unadjusted analyses, or studies with different approaches to missing data can yield different estimates.
This does not mean that you should automatically accept whichever result comes from the design you consider superior. Instead, examine whether methodological differences provide a plausible explanation for the pattern and whether some studies should consequently contribute more heavily to the overall interpretation.
Context may be part of the effect rather than background noise
Studies conducted in different countries, institutions, healthcare systems, classrooms, labor markets, or policy environments may produce different findings because the phenomenon operates differently under those conditions.
For example, an educational intervention that depends heavily on reliable internet access may perform differently across settings with different technological infrastructure. Treating context as irrelevant could erase an important boundary condition of the intervention.
In such cases, disagreement can improve the synthesis. Instead of asking only “Does it work?” you may be able to ask the more informative question, “Under what conditions does it appear to work?”
Chance and imprecision can also produce different-looking results
Not every disagreement has a substantive explanation. Small studies can produce unstable estimates, and sampling variation can make effects appear different even when the underlying effects are similar.
Confidence intervals and other measures of uncertainty can help determine whether apparently different estimates are actually compatible with a common range of plausible effects.
Explanations discovered after seeing the results require caution
Once conflicting findings are visible, it is easy to search study characteristics until something appears to explain them. Perhaps positive studies used one age group, one country, one instrument, or one implementation model.
That pattern may be meaningful. It may also be coincidental.
Cochrane guidance therefore distinguishes more credible pre-specified investigations of heterogeneity from post hoc explorations generated after the study results are known. Post hoc findings can generate hypotheses, but they generally should not be presented with the same confidence as an explanation specified in advance and supported by appropriate analysis.
Observed disagreement
A defensible description that relevant study findings differ in a way that matters for interpretation.
Explanation for disagreement
A separate inference about why those findings differ, which requires its own evidence and may remain uncertain.
Sometimes the correct explanation is that you do not know
A literature synthesis does not need to solve every contradiction. If studies disagree and the available evidence cannot distinguish among plausible explanations, say so.
Unexplained inconsistency can itself reduce confidence in an overall conclusion. Formal frameworks such as GRADE explicitly treat important unexplained inconsistency as a reason for lower certainty in a body of evidence.
Preserving that uncertainty is preferable to supplying an attractive explanation that the evidence cannot actually establish.
04 · A Practical Example
When Conflicting Findings Reveal a More Useful Question
Hypothetical Example
Does automated feedback improve student writing?
Imagine six hypothetical studies. Three report meaningful improvements in writing performance, two find little or no difference, and one reports improvement only among students with higher baseline writing proficiency.
Tempting synthesis
“Most studies show that automated feedback improves student writing.”
Inspect the disagreement
The positive studies use repeated feedback over an entire semester. The two null studies involve a single feedback session. The remaining study reports different results according to baseline proficiency.
Evaluate the explanation
Duration of exposure and baseline proficiency are plausible explanations, but whether they were specified in advance and whether enough evidence exists to support those explanations must be considered.
More defensible synthesis
“Findings are mixed. Studies using repeated feedback tend to report larger improvements than single-session studies, while one study suggests that effects may also vary by baseline proficiency. These patterns may help explain the disagreement, but the available evidence is insufficient to establish either factor as the cause of the variation.”
The revised synthesis is longer, but it is also more informative. It distinguishes the observed evidence from the proposed explanation and shows readers where uncertainty remains.
07 · A Quick Checklist
Have You Represented Conflicting Evidence Fairly?
When important studies disagree, check:
I have established that the findings genuinely differ rather than relying on differences in statistical significance alone.
I have not hidden conflicting studies behind a majority count or an overly simple overall statement.
I have compared relevant differences in populations, interventions or exposures, outcomes, methods, implementation, follow-up, and settings.
I have considered whether differences in study quality or evidential relevance may help explain the pattern.
I distinguish observed disagreement from my explanation for why the disagreement occurs.
I identify post hoc explanations as exploratory rather than presenting them as established causes.
If the disagreement remains unexplained, my conclusion preserves that uncertainty.