Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Questions Remain Unanswered Because Evidence Is Inconsistent?

Conflicting findings do not simply mean that researchers need one more study. Learn how to identify the question hidden inside inconsistent evidence and investigate why results differ.

651
Unanswered Questions From Inconsistent Evidence Guide 651 of 899
01 · The Question

What Should You Ask When Studies Seem to Disagree?

One study reports a substantial benefit. Another finds almost no difference. A third reports an effect only for certain participants. A fourth points in the opposite direction.

The familiar conclusion is that “findings are mixed” and more research is needed.

That diagnosis is usually incomplete.

Inconsistent evidence may indicate random variation, methodological differences, bias, differences in populations or interventions, measurement choices, contexts, analytical decisions, or genuine variation in how a phenomenon operates.

The important research question is often not simply, “Which study is correct?” It is, “Why do credible estimates differ, and under what conditions should we expect different answers?”

02 · The Short Answer

Treat Inconsistency as Something to Explain, Not Merely Average Away

In Brief

Questions remain unanswered because evidence is inconsistent when studies provide meaningfully different estimates or conclusions and the available evidence cannot adequately explain why those differences occur.

The next research need is not automatically another study estimating the same overall effect. First determine whether the apparent disagreement reflects chance, differences in methods or populations, bias, outcome definitions, intervention implementation, or genuine effect heterogeneity.

03 · What You Need to Know

Find Out What the Disagreement Is Trying to Tell You

Inconsistency is more than one significant result and one non-significant result

A common mistake is to classify studies according to whether their individual p-values cross a significance threshold and then declare the literature inconsistent.

That can be misleading.

Suppose one study estimates an effect of 0.30 with a narrow enough confidence interval to produce p <.05, while another estimates 0.28 with a wider confidence interval and p >.05. The studies may actually provide very similar estimates.

The apparent disagreement comes from differences in precision, not necessarily differences in the underlying effect.

Cochrane's guidance on assessing inconsistency instead emphasizes differences in effect estimates, confidence-interval overlap, and measures of between-study heterogeneity. It explicitly advises against reducing interpretation to statistical significance alone.

Different statistical significance Studies fall on different sides of a threshold such as p <.05.
Different effects Studies estimate effects that differ meaningfully in magnitude, direction, or practical interpretation.

The second is the more important starting point for diagnosing inconsistency.

Some variation is expected even when studies estimate the same underlying effect

Samples vary. Measurements contain error. Estimates fluctuate.

Two well-designed studies therefore need not produce identical numbers. Observed differences may be compatible with sampling variation rather than genuine differences in the underlying phenomenon.

The challenge is deciding whether variation is larger or more consequential than would reasonably be expected from sampling error alone.

Meta-analysis provides statistical tools for examining heterogeneity, but no single statistic settles the matter. Measures such as I² and estimates of between-study variance can help characterize heterogeneity, while the actual direction, magnitude, precision, and clinical or practical implications of study estimates remain important. Cochrane recommends examining several indicators rather than treating a single heterogeneity statistic mechanically.

Real inconsistency may reveal effect modification

Sometimes studies disagree because the effect genuinely differs across conditions.

An intervention may work better for beginners than experts. A policy may have different effects where implementation capacity is high. An educational technology may benefit students with strong digital access while producing little advantage where access is unreliable.

If these differences are credible, the question “Does it work?” may be too broad.

A more useful question becomes: for whom, under what conditions, compared with what, and for which outcomes does it work?

GRADE treats unexplained inconsistency as a reason for lower certainty in a body of evidence. When robust explanations identify meaningful modifiers, however, the variation may be scientifically informative rather than merely inconvenient.

Population differences may explain apparently conflicting findings

Studies may recruit participants with different baseline characteristics, levels of prior experience, risks, needs, ages, socioeconomic circumstances, or other relevant characteristics.

If those characteristics modify the effect, pooled conclusions can hide important variation.

This is where inconsistency can point toward a population question. Rather than simply declaring that findings conflict, investigate whether population characteristics change the applicability or magnitude of the finding.

The intervention may not actually be the same across studies

Two papers can use the same label for substantially different interventions.

“Online learning,” “AI-assisted instruction,” “peer feedback,” “mindfulness,” or “professional development” can describe families of interventions whose intensity, content, duration, implementation, support, and fidelity differ considerably.

If intervention versions differ, heterogeneous results may be unsurprising.

The unanswered question may therefore concern which component, dose, implementation strategy, or version produces the effect.

Different outcomes can manufacture apparent disagreement

Suppose one study reports improved engagement, another finds no difference in examination scores, and a third reports increased workload.

Those findings do not necessarily contradict one another. They concern different consequences.

Likewise, two studies may both claim to measure “achievement” while one uses a standardized assessment and the other course grades. Before diagnosing inconsistency, determine whether the studies are actually estimating comparable outcomes.

If not, the more fundamental problem may be that researchers measure different or poorly aligned outcomes.

Design differences can produce different answers

Randomized trials, observational studies, cross-sectional surveys, longitudinal cohorts, quasi-experiments, and uncontrolled pre-post studies do not necessarily estimate the same quantity or protect equally against the same biases.

If weaker designs consistently produce larger effects while stronger designs produce smaller ones, that pattern itself deserves investigation.

The research gap may therefore involve designs that cannot adequately support the intended inference rather than a mysterious disagreement about the phenomenon.

Bias can create apparent heterogeneity

Differences across studies may also reflect varying risks of bias. Selective reporting, missing data, confounding, measurement problems, deviations from intended interventions, and other methodological issues can move estimates in different directions.

This is one reason inconsistency should not immediately be celebrated as evidence of interesting contextual variation. Before theorizing about moderators, examine whether some of the disagreement can be explained by study credibility.

There are several different forms of inconsistency

Pattern in the literature Possible interpretation Question worth investigating
Effects point in opposite directions Potential genuine effect differences, bias, or major methodological variation What conditions determine the direction of the effect?
Effects point in the same direction but vary greatly in magnitude Possible effect modification or differences in implementation What determines how large the effect becomes?
Point estimates are similar but significance differs Different precision rather than substantive inconsistency Is there actually meaningful heterogeneity to explain?
Different populations produce different effects Possible population-level effect modification Which characteristics modify the effect?
Different intervention versions produce different effects Possible variation in dose, fidelity, components, or implementation Which intervention features matter?
Different designs produce systematically different estimates Possible bias or differences in what is being estimated How does methodological design influence the apparent effect?

A pooled average can be correct and still conceal the important question

Meta-analysis can summarize evidence by estimating an average effect under appropriate assumptions. But an average does not necessarily describe every setting or participant.

Imagine that an intervention produces substantial benefits in one context and little benefit in another. The average effect may be estimated accurately while still being a poor answer to the practical question facing someone in either context.

If heterogeneity is consequential, understanding its sources can be more informative than simply obtaining a more precise average.

Watch Out

Do not automatically interpret a statistically significant heterogeneity test or a high I² value as proof that a particular moderator explains the variation. Identifying heterogeneity and explaining heterogeneity are separate tasks, and subgroup explanations require appropriate evidence.

Subgroup hunting can create false explanations

Once researchers observe inconsistency, it is tempting to divide studies repeatedly until some characteristic appears to explain it.

This can generate attractive stories from chance patterns.

Potential moderators are more convincing when specified from theory or prior evidence, measured adequately, supported by enough studies or participants, and examined using appropriate interaction or meta-regression methods. Even then, observational comparisons across studies can be confounded.

“The effect differs between these subgroups” requires stronger evidence than noticing that one subgroup has a statistically significant effect and another does not.

Sometimes inconsistent evidence is itself the research opportunity

AHRQ's framework for determining research gaps explicitly identifies inconsistency or unknown consistency as one reason existing evidence may fail to support a conclusion. The framework emphasizes identifying not merely where evidence is missing but why it falls short.

That distinction is useful. A literature containing twenty conflicting studies may present a more consequential gap than a literature containing no study of an arbitrary new variable combination.

The contribution comes from explaining the uncertainty, not from increasing the study count to twenty-one.

04 · A Practical Example

When “Mixed Findings” Become a Better Research Question

Hypothetical Example

Does generative AI improve university students' academic writing?

Imagine a hypothetical literature containing 24 studies. Eight report substantial improvements, nine report modest improvements, five find little difference, and two report worse independent writing performance among frequent AI users.

Weak interpretation The literature is inconsistent, so another study is needed to determine whether AI improves writing.
Evidence audit The studies differ in whether students use AI for brainstorming, feedback, rewriting, or complete text generation. They also differ in prior writing ability, duration of use, outcome measurement, and whether writing is assessed with or without AI assistance.
Better uncertainty The effect may depend on how AI is used, which writing outcome is assessed, and whether the outcome represents assisted performance or independent skill.
Research implication A useful study would test a credible explanation for the variation rather than merely generate another overall estimate under yet another set of conditions.

The conflicting literature has not disappeared. It has become more informative. Instead of asking only whether AI “works,” the researcher can investigate why effects differ and which conditions produce the outcomes of interest.

05 · What Researchers Often Get Wrong

Common Mistakes When Interpreting Conflicting Evidence

Misconception

One Significant Study and One Non-Significant Study Contradict Each Other

Not necessarily. Their effect estimates may be nearly identical while their precision differs. Compare estimates and uncertainty directly rather than classifying studies solely according to whether individual p-values cross a significance threshold.

Misconception

The Largest Study Tells Us Which Result Is Correct

Larger studies usually provide greater precision, but sample size does not automatically eliminate bias, measurement problems, indirectness, or design limitations. Study credibility and relevance still matter.

Misconception

A Meta-Analysis Eliminates Inconsistency

Pooling studies can estimate an average effect, but it does not make genuine heterogeneity disappear. When effects differ meaningfully across settings or populations, the reasons for that variation may remain central to interpretation.

Misconception

A High I² Tells You Why Studies Differ

No. I² describes one aspect of observed heterogeneity relative to sampling error; it does not identify its cause. Explanations require substantive and methodological investigation.

Misconception

Finding a Significant Effect in One Subgroup but Not Another Proves a Subgroup Difference

No. Evidence that one subgroup differs from another requires an appropriate comparison of effects, such as an interaction analysis, rather than separate significance tests within each subgroup.

Misconception

The Solution to Conflicting Evidence Is Always Another Replication

Replication can be valuable, but repeating the same design under similar conditions may simply add another estimate to the disagreement. When plausible explanations for heterogeneity exist, research designed to distinguish among those explanations may be more informative.

06 · What This Means for You

Turn “Mixed Findings” Into Testable Explanations

If your literature review ends with “previous studies produced mixed results,” you have identified a symptom. The next step is diagnosis.

Map the studies according to characteristics that could plausibly explain the differences: populations, interventions or exposures, comparisons, outcomes, time points, settings, designs, measurement methods, implementation, and risk of bias.

A simple decision framework

If effect estimates are actually similar
Do not manufacture inconsistency from different p-values or verbal conclusions.
If estimates differ beyond plausible sampling variation
Identify theoretically and methodologically credible sources of heterogeneity.
If population differences align with effects
Test whether relevant population characteristics genuinely modify the effect.
If intervention or exposure definitions vary
Determine whether different versions, doses, components, or implementation conditions account for different outcomes.
If methodological quality aligns with effect size
Investigate whether bias or design differences explain part of the apparent inconsistency.
If no credible explanation emerges
Treat unexplained heterogeneity as uncertainty rather than inventing a post hoc story.

The goal is not to force every difference into a neat explanation. Sometimes the evidence genuinely does not yet reveal why estimates vary. Recognizing that uncertainty is more defensible than selecting whichever explanation makes the proposed study easiest to justify.

07 · A Quick Checklist

Before Claiming That the Literature Is Inconsistent, Check the Evidence

Before proposing research based on conflicting findings, check:
Have I compared effect estimates and their uncertainty rather than only p-values or authors' verbal conclusions?
Are the studies actually estimating sufficiently similar outcomes and effects to be considered contradictory?
Could observed variation reasonably reflect sampling error rather than genuine effect differences?
Do populations, settings, interventions, exposures, comparisons, outcomes, or follow-up periods differ in ways that could modify the result?
Do differences in study design, measurement, implementation, or risk of bias align with differences in findings?
Are proposed moderators supported by theory or prior evidence rather than discovered only after inspecting the results?
Would the proposed study test an explanation for inconsistency rather than simply add another estimate?
If inconsistency cannot currently be explained, have I preserved that uncertainty rather than forcing a conclusion?
08 · Frequently Asked Questions

Questions About Inconsistent Research Evidence

What does inconsistent evidence mean?

It generally means that relevant studies provide meaningfully different estimates or conclusions that cannot yet be adequately reconciled. Some variation is expected from sampling error, so not every numerical difference constitutes important inconsistency.

Do different p-values mean that two studies conflict?

No. One study can produce p <.05 and another p >.05 even when their estimated effects are similar. Compare the effects, confidence intervals, precision, and substantive interpretations rather than significance labels alone.

What is heterogeneity in research?

Heterogeneity refers to variation among studies. It can involve differences in participants, interventions, exposures, outcomes, methods, settings, or estimated effects. Statistical heterogeneity specifically concerns variation in effect estimates beyond what may be expected from sampling error.

Does a high I² mean a meta-analysis is invalid?

Not automatically. I² is one indicator of heterogeneity and should be interpreted alongside the magnitude and direction of effects, between-study variance, confidence intervals, study characteristics, and substantive context. High heterogeneity may make a single pooled estimate less informative for some questions, but it does not mechanically invalidate synthesis.

How can I identify why studies disagree?

Compare studies systematically across plausible effect modifiers and methodological characteristics, including populations, interventions, comparisons, outcomes, settings, time points, designs, measurements, implementation, and risk of bias. Explanations should ideally be grounded in theory or prior evidence rather than generated solely from observed patterns.

Should I conduct another study if previous findings are mixed?

Only if the new study can add information that helps resolve the uncertainty. This might involve testing a credible moderator, improving the design, standardizing an important outcome, studying a relevant population, or obtaining more precise evidence. Simply repeating the same limitations may not clarify the disagreement.

Can inconsistent evidence itself be an important finding?

Yes. Genuine heterogeneity can reveal that an effect is conditional rather than universal. Understanding when, where, for whom, or under what implementation conditions an effect changes may be more informative than estimating a single average effect.

09 · The Bottom Line

Do Not Stop at Saying the Findings Are Mixed

The Bottom Line

Inconsistent evidence creates an important unanswered question when studies provide meaningfully different estimates and the available evidence cannot yet explain why those differences occur.

Compare effects rather than significance labels, investigate credible sources of heterogeneity, and distinguish genuine effect modification from methodological artifacts and chance. The most informative next study may not ask once again whether an effect exists, but why its magnitude or direction changes across conditions.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes