01 · The Question
What should you do when credible studies point in different directions?
One study reports a substantial effect. Another finds little difference. A third finds the relationship only in certain participants. A fourth reaches the opposite conclusion.
It is tempting to resolve this by choosing the study with the largest sample, newest publication date, strongest reputation, or result closest to your own expectation. That may make the literature review easier to write, but it does not explain the evidence.
Disagreement can arise because studies differ in populations, interventions, exposures, outcomes, designs, measurements, implementation, analysis, bias, or random sampling variation. In some cases, the phenomenon itself genuinely varies across contexts. Cochrane distinguishes clinical diversity, methodological diversity, and statistical heterogeneity when examining variation among studies.
The important question is therefore not simply which study should you believe, but why do the results differ and what does that difference tell you?
03 · What You Need to Know
What can cause research findings to disagree?
First establish that there is a disagreement worth explaining
Two point estimates do not need to be identical for studies to tell broadly compatible stories. Every estimate is subject to sampling uncertainty, and estimates from separate samples will vary even when the underlying effect is similar.
Suppose one study estimates an effect of 0.25 and another estimates 0.34. Calling those findings contradictory merely because the numbers differ would usually be difficult to justify. Their uncertainty intervals, direction, and substantive implications may overlap considerably.
Conversely, two estimates can both be described as “statistically significant” while differing substantially in magnitude. Significance labels are therefore poor tools for deciding whether studies agree.
Compare the estimates, their uncertainty, direction, and substantive meaning rather than simply comparing p-values.
Some disagreement is expected from chance alone
Separate studies sample different participants and therefore produce different estimates. This sampling variation means perfectly identical results should not be expected even when studies estimate the same underlying quantity.
In meta-analysis, statistical heterogeneity refers to variability in effect estimates beyond what would be expected from sampling error alone. Cochrane recommends assessing the presence and extent of between-study variation rather than assuming that every observed difference represents a genuine difference in effects.
Sampling variation
Estimates differ because separate samples will not produce exactly the same numerical result even when they estimate the same underlying quantity.
Heterogeneity
Effects or estimates vary across studies beyond what would reasonably be attributed to sampling variation alone.
The effect itself may genuinely differ across populations
A phenomenon does not have to behave identically for everyone.
An educational intervention may work differently for novice and advanced learners. A behavioral association may differ across age groups. An intervention requiring substantial digital infrastructure may perform differently in institutions with very different technological resources.
These are examples of effect modification: the effect changes according to another characteristic.
Cochrane notes that clinical diversity in participant characteristics or interventions can produce heterogeneity when those characteristics modify the intervention effect.
If credible effect modification exists, disagreement can be informative. The better conclusion may not be “the intervention works” or “the intervention does not work,” but “its effect appears to depend on particular conditions.”
The intervention or exposure may not actually be the same
Research labels can conceal substantial variation.
Two papers may both investigate “AI-assisted learning,” while one provides students with automated hints during practice and another allows unrestricted use of a chatbot for assignment completion. Studies of “peer feedback,” “online learning,” “mindfulness,” or “active learning” can similarly implement quite different interventions under the same broad label.
Dose, duration, intensity, implementation quality, adherence, timing, and comparison conditions can all matter.
Before explaining conflicting results statistically, ask whether participants actually received comparable experiences.
Different outcomes can produce legitimately different conclusions
An intervention can improve one outcome while doing little for another.
For example, an educational technology might improve immediate task performance without improving long-term retention. It might increase student satisfaction while having little effect on achievement. It could reduce completion time while increasing error rates.
Those findings are not contradictory. They concern different consequences.
Cochrane emphasizes that populations, interventions, comparators, and outcomes define the questions being synthesized and should be considered when determining how studies are grouped.
Measurement choices can create apparent disagreement
Even when researchers claim to measure the same construct, their instruments may operationalize it differently.
One study might measure critical thinking using a validated performance assessment. Another may ask students to rate their own critical-thinking ability. A third might infer critical thinking from course grades.
The label is shared. The measurement is not.
Measurement differences can change both what is being estimated and the amount of bias affecting the estimate. JBI's appraisal guidance explicitly treats valid and reliable outcome measurement as a core consideration when evaluating observational evidence.
Study design can change what the result means
A randomized experiment, prospective cohort, cross-sectional survey, and qualitative study can all investigate aspects of the same broad topic without producing directly interchangeable evidence.
Suppose a cross-sectional survey finds that students who use AI more frequently have lower grades. A randomized experiment finds that providing AI-assisted feedback improves performance on a writing task. These results might initially sound contradictory.
They need not be. Students who voluntarily use AI frequently may differ systematically from students who use it less. The randomized experiment estimates something different: the effect of assigning a particular AI-supported activity under specified conditions.
Design-specific sources of bias also matter. JBI highlights concerns such as exposure classification, confounding, temporal precedence, outcome measurement, participant retention, and statistical conclusion validity when appraising cohort studies.
Confounding can make observational studies disagree
Two observational studies may adjust for different sets of variables. One may measure important confounders well while another omits them. Even when both use regression adjustment, they may not be estimating equivalent quantities.
Residual confounding can remain after adjustment, particularly when confounders are poorly measured or omitted. Cochrane specifically identifies residual confounding and biases that vary across non-randomized studies as potential sources of heterogeneity.
Therefore, disagreement among adjusted estimates should not be interpreted solely from the final coefficients. Examine what was measured and what the models actually controlled.
Analytical decisions can change the result
Researchers make many analytical choices: how variables are coded, which covariates are included, how missing data are handled, whether outliers are excluded, which statistical model is used, how subgroups are defined, and which outcome time point is emphasized.
Reasonable choices can sometimes produce meaningfully different estimates. Poor choices can introduce additional bias.
If two studies use the same or similar data but reach different conclusions, analytical specifications deserve particular scrutiny.
Bias can push different studies in different directions
Heterogeneity is not always evidence of interesting real-world variation. Methodological weaknesses can also produce different estimates.
Cochrane notes that differences in design, outcome measurement, and risk of bias can produce methodological diversity and corresponding heterogeneity. When heterogeneity arises from differing degrees of bias, the studies may not be estimating the same quantity reliably.
This is why understanding disagreement requires critical evaluation of the studies, not simply comparison of their conclusions.
Do not treat I² as a disagreement detector with a universal cutoff
In meta-analysis, the I² statistic is commonly used to describe the proportion of observed variability in effect estimates associated with heterogeneity rather than sampling error. But Cochrane warns that thresholds for interpreting I² can be misleading because its importance depends on factors including the magnitude and direction of effects and the strength of evidence for heterogeneity.
A numerical heterogeneity statistic can help describe variation. It cannot tell you why the studies differ.
Watch Out
Do not convert a heterogeneity statistic into an explanation. “I² was high” describes a pattern in a particular meta-analysis. It does not tell you whether the cause is population differences, intervention variation, measurement, bias, analysis, or something else.
Subgroup explanations are easy to invent after seeing the results
Once disagreement is visible, almost any study characteristic can become a tempting explanation. Perhaps studies in one country look different. Perhaps younger participants respond differently. Perhaps longer interventions work better.
Some explanations will be real. Others will be patterns produced by chance.
Cochrane advises that subgroup analyses and meta-regression used to explore heterogeneity require caution, particularly when explanations are devised after results are known. Post hoc explorations are generally better treated as hypothesis-generating than as definitive explanations.
The more explanations you try, the easier it becomes to find one that appears to fit.
Sometimes the correct conclusion is that the disagreement remains unexplained
Researchers understandably want resolution. But evidence does not always provide one.
After examining populations, interventions, outcomes, measurements, design, bias, and analysis, meaningful differences may remain. GRADE treats important unexplained inconsistency as a reason for lower certainty in a body of evidence.
That is not a failure of synthesis. It is a substantive finding: the current literature does not yet support one stable general estimate or explanation.