01 · The Question
What Should You Do When Systematic Reviews Disagree?
You search for reviews addressing your question and find something more complicated than an empty literature: several reviews already exist, but they do not agree.
One concludes that an intervention is effective. Another reports insufficient evidence. A third reaches a more cautious conclusion or produces a noticeably different pooled estimate.
That looks like an obvious review gap. If the reviews disagree, perhaps another systematic review can settle the issue.
Sometimes it can. But conflicting conclusions can arise for many reasons, including differences in the questions asked, studies included, search dates, risk-of-bias judgments, analytical decisions, and interpretation. Before conducting another review, you need to understand the source of the discordance. Otherwise, you may simply add one more conclusion to an already crowded argument.
03 · What You Need to Know
Why Systematic Reviews Can Reach Different Conclusions
First determine whether the reviews are genuinely conflicting
Two reviews can reach different conclusions without actually contradicting each other.
Suppose one review asks whether an intervention improves short-term performance among university students, while another examines long-term outcomes across adolescents and adults. One reports a beneficial effect and the other finds uncertain evidence. The conclusions sound different, but the reviews are not necessarily answering the same question.
Before diagnosing conflict, establish whether the reviews actually address the same or sufficiently similar research question.
Apparent disagreement
Reviews reach different conclusions because they examine meaningfully different questions, populations, interventions, outcomes, or contexts.
Substantive disagreement
Reviews address sufficiently similar questions but produce materially different findings or interpretations that cannot be explained simply by differences in scope.
The second situation creates a stronger reason to investigate whether another synthesis could add value.
Compare which primary studies entered each review
Reviews addressing the same broad question may contain different primary studies.
One review may have searched more databases. Another may have searched later. Eligibility criteria may differ subtly. Language or publication-status restrictions may exclude studies. One review may include non-randomized evidence while another restricts inclusion to randomized trials.
Cochrane guidance for overviews recognizes that overlapping reviews frequently contain many, but not all, of the same primary studies. It recommends mapping which studies occur in which reviews because differences and overlap in the underlying evidence can materially affect the synthesis.
This comparison is often one of the fastest ways to understand discordance. If Review A includes ten studies and Review B includes six, with only four studies shared, you are not really comparing two analyses of exactly the same evidence.
Search dates can produce different answers
Two reviews may have similar eligibility criteria but different evidence cutoffs.
If one searched through 2019 and another through 2025, the later review may contain studies unavailable to the earlier authors. Its different conclusion may therefore reflect an evolving evidence base rather than methodological disagreement.
In that situation, another review is not automatically necessary. The more recent review may already provide the needed update, assuming its search and methods are adequate. Conversely, a recent review with a poor search may still omit consequential evidence.
Eligibility criteria can create different evidence bases
Small differences in eligibility rules can have large consequences.
One review may include only a narrowly defined intervention, while another combines several related interventions. One may exclude studies below a minimum duration. Another may include particular outcome measures that the first excludes.
These choices are not necessarily errors. Review questions require boundaries. The important issue is whether the boundaries explain the conflicting findings and which formulation is appropriate for the question you need answered.
Different risk-of-bias judgments can change interpretation
Two reviews may include nearly identical studies yet evaluate their credibility differently.
Risk-of-bias assessments involve structured judgments. Differences can arise because reviewers use different tools, apply criteria differently, or evaluate different outcomes and results within the same study.
If one review treats most evidence as credible while another identifies serious limitations, similar effect estimates can lead to different conclusions about how confidently those estimates should be interpreted.
This distinction matters because disagreement may reside less in the numerical result than in the confidence placed in it.
Analytical decisions can produce different pooled estimates
Meta-analysis is not a single automatic calculation. Reviewers make decisions about effect measures, statistical models, handling of multiple outcomes, missing data, cluster designs, time points, heterogeneity, and other analytical issues.
Cochrane guidance emphasizes assessing between-study variation when conducting meta-analysis and recognizes that analytical decisions can affect estimates and their interpretation. Sensitivity analyses can be useful for assessing whether findings are robust to reasonable methodological choices.
If conflicting reviews differ mainly because of analytical choices, another review may be valuable when it can transparently evaluate those choices and test the robustness of the conclusions rather than merely select a preferred model.
The numbers may agree more than the words do
Sometimes reviews report similar effect estimates yet write substantially different conclusions.
One author group might describe an effect as “promising,” another as “small,” and another as “insufficient to support implementation.” Differences in certainty judgments, thresholds for practical importance, or emphasis on harms and limitations can produce verbal discordance without major statistical disagreement.
Compare the actual estimates, confidence intervals, certainty assessments, and evidence limitations before assuming that contradictory prose means contradictory evidence.
Watch Out
Do not classify reviews as conflicting solely because one uses positive language and another uses cautious language. Compare the underlying results and certainty of evidence before deciding that a substantive contradiction exists.
Heterogeneity may be the real unresolved question
Different studies may genuinely produce different effects. If so, asking for a single universal answer may be part of the problem.
Differences in populations, interventions, implementation, settings, follow-up periods, or study methods can produce heterogeneity. Cochrane recommends investigating plausible sources of heterogeneity where appropriate, while cautioning that subgroup analyses and meta-regression can themselves be misleading when poorly planned or interpreted.
A useful new review might therefore ask not merely “Does it work?” but “Under what conditions do the results differ, and is there credible evidence explaining that variation?”
Another review should resolve something, not cast another vote
When two reviews disagree, it is tempting to conduct a third and see which side it supports. That is not a strong scientific rationale.
The more useful objective is to identify why the reviews disagree and design the new synthesis around that unresolved problem. Perhaps one search omitted important studies. Perhaps methodological quality differs sharply. Perhaps newer evidence changes the result. Perhaps different population boundaries explain the divergence.
PRISMA 2020 requires authors to describe the rationale for a review in the context of existing knowledge. Its explanation and elaboration guidance identifies discordant results among existing reviews as one situation that may warrant explaining why another review is necessary.
That does not mean disagreement automatically authorizes another review. It means disagreement can form part of a defensible rationale when the new review has a credible way to clarify it.
Sometimes an overview is more appropriate than another review of primary studies
If several systematic reviews already address closely related questions, your real research problem may concern the reviews themselves.
An overview of reviews can systematically examine multiple reviews, their overlap, methodological quality, results, and conclusions. Cochrane notes that overviews can help address situations involving overlapping reviews with variable conduct, quality, reporting, results, or conclusions.
Whether an overview is appropriate depends on your objective. If the unresolved problem requires re-examining primary studies or conducting a new synthesis from the original evidence, another systematic review may be more suitable. If the purpose is to understand why existing syntheses disagree, review-level analysis may sometimes answer the question more directly.
Several conflicting reviews can still leave a useful gap
Conflict is one example of how several existing reviews can coexist with a meaningful review gap. The key is to define that gap precisely.
“Previous reviews disagree” is an observation. “Previous reviews disagree because they include different evidence and use incompatible analytical decisions, and no synthesis has tested whether the conclusion is robust to those differences” is the beginning of a research rationale.
07 · A Quick Checklist
Check Whether Conflicting Reviews Justify Another Synthesis
Before starting another review, check:
Confirm that the reviews address sufficiently similar questions to make their conclusions genuinely comparable.
Compare populations, interventions or exposures, comparators, outcomes, settings, and eligible study designs.
Map which primary studies are shared and which appear in only one review.
Compare databases, search strategies, restrictions, and final search dates.
Compare risk-of-bias assessments and other judgments about evidence credibility.
Compare effect measures, statistical models, handling of heterogeneity, sensitivity analyses, and other consequential analytical choices.
Determine whether the numerical results actually conflict or whether the main difference lies in interpretation.
State precisely how another review would explain or reduce the unresolved disagreement.