01 · The Question
What Has a Study Actually Shown When the Difference Is Not Significant?
You compare two groups and obtain p =.18. The result is not statistically significant. Can you now write that there was “no difference” between the groups?
Usually, not on that basis alone. A nonsignificant result may occur because the true difference is negligible. It can also occur because the data are too imprecise to distinguish a meaningful difference from random variation.
The distinction matters because “we did not establish a difference” and “we established that the groups do not meaningfully differ” are different claims. Conventional significance testing answers the first question much more directly than the second.
02 · The Short Answer
A Nonsignificant Test Does Not Establish Equality
In Brief
“No significant difference” fails to show that there is no difference whenever the data remain compatible with differences large enough to matter. A conventional nonsignificant result tells you that the analysis did not reject the null hypothesis at the chosen threshold; it does not demonstrate that the null hypothesis is true.
To judge whether a meaningful difference has been ruled out, examine the estimated effect and its uncertainty, consider the study's ability to provide an informative estimate, and, when the research question concerns similarity or practical equivalence, use an inferential approach designed to address that question.
03 · What You Need to Know
Why Statistical Nonsignificance and No Difference Are Not Equivalent
Conventional significance testing is asymmetric
In a conventional null-hypothesis significance test, researchers typically begin with a null hypothesis such as a population difference of zero. They then ask whether the observed data would be sufficiently unusual under that hypothesis to reject it at a chosen significance level.
If p falls below the threshold, researchers reject the null hypothesis according to that testing procedure. If p does not, they fail to reject it.
Those outcomes are deliberately asymmetric. Failing to reject a hypothesis is not the same inferential act as demonstrating that the hypothesis is correct. As Altman and Bland pointed out in their classic discussion of nonsignificant findings, a study that does not find a statistically significant difference has often shown only an absence of evidence for a difference rather than evidence that the difference is absent .
A p-value does not tell you how large the difference might still be
Suppose two studies both report p =.12. The identical significance classification does not mean the studies contain the same information.
One might estimate a mean difference of 4 points with a 95% confidence interval from -2 to 10. Another might estimate a difference of 0.3 points with an interval from -0.5 to 1.1. Both results can be nonsignificant, yet the first remains compatible with a fairly substantial positive difference while the second is much more tightly concentrated around small differences.
The statement “not significant” suppresses this distinction. It tells you about a threshold decision under a particular hypothesis test, not the complete range of effect sizes still compatible with the evidence.
Wide uncertainty can make “no difference” indefensible
Confidence intervals help reveal what a binary significance label hides. If an interval includes zero but also contains effects that would matter scientifically or practically, the data have not resolved whether an important difference exists.
This is particularly common in small or noisy studies. A study can fail to detect a real difference because its estimate is imprecise.
In that situation, the appropriate interpretation is uncertainty. The study has not established the difference, but neither has it provided strong evidence that the difference is negligible.
No statistically significant difference
The chosen test did not reject the null hypothesis at the specified significance level.
No meaningful difference
Differences large enough to matter have been excluded or otherwise made sufficiently implausible by an analysis capable of addressing that question.
Statistical power affects how surprising a null result should be
A study with little ability to detect effects of the size researchers care about can readily produce a nonsignificant result even when such an effect exists. Consequently, a null finding from an underpowered study may carry little information about whether the underlying effect is negligible.
This is why simply accumulating many underpowered studies with nonsignificant results does not automatically establish that nothing happens.
Power calculations are particularly useful during study planning because they force researchers to consider what effect sizes a design is intended to detect. After data have been observed, interpretation should not stop at a retrospective declaration that a study was “powered” or “underpowered.” The observed effect estimate, its uncertainty, and the design assumptions provide more direct information about what the study has and has not ruled out.
Similarity requires defining what counts as similar
Two populations will rarely be exactly identical on a continuously measured outcome. With sufficiently large samples, even tiny differences may become statistically detectable. Conversely, a small study may fail to detect a substantial difference.
The scientifically useful question is therefore often not “Is the true difference exactly zero?” but “Is the difference small enough to be considered negligible for this purpose?”
Answering that requires a substantive threshold. In equivalence testing, researchers specify lower and upper equivalence bounds representing differences considered too small to matter. The two one-sided tests procedure, commonly abbreviated TOST, then tests whether effects at or beyond those bounds can be rejected.
This directly addresses a different question from a conventional significance test and can provide meaningful evidence that an effect is absent or sufficiently small .
Four outcomes are possible when significance and equivalence are considered together
Once researchers distinguish testing for a difference from testing whether meaningful differences can be excluded, a useful feature becomes visible: “significant” and “equivalent” are not simple opposites.
Result
What it can suggest
What it does not establish
Significant, not equivalent
Evidence of a difference, with meaningful effects not excluded
That the difference is necessarily important
Nonsignificant, equivalent
No conventional evidence of a difference, while effects beyond the equivalence bounds are rejected
That the true difference is exactly zero
Significant, equivalent
A statistically detectable difference that nevertheless falls within justified equivalence bounds
That statistical detectability makes the difference substantively important
Nonsignificant, not equivalent
The data are insufficient to establish either a conventional difference or equivalence
That the groups are the same
The last situation is especially important. Researchers sometimes interpret it as evidence for “no difference,” when the study may simply be unable to discriminate between a negligible effect and an important one.
04 · A Practical Example
Two Studies Can Both Be Nonsignificant but Support Different Conclusions
Hypothetical Example
Comparing two teaching approaches
Suppose researchers compare examination scores under Teaching Method A and Teaching Method B. Before collecting data, they justify differences smaller than 5 percentage points in either direction as educationally negligible for the decision they need to make.
Study A
The estimated difference is 2 points, with a 95% confidence interval from -8 to 12 points. The conventional test is nonsignificant.
What Study A tells you
The interval includes zero, but it also includes differences exceeding the ±5-point threshold in both directions. The study has not established a difference, yet it has not established practical similarity either.
Study B
The estimated difference is 0.5 points, with a 95% confidence interval from -2 to 3 points. The conventional test is also nonsignificant.
What Study B tells you
Under the hypothetical ±5-point criterion, the evidence is much more informative because the interval is concentrated within differences already defined as educationally negligible. An appropriate equivalence analysis could formally test the corresponding equivalence claim.
Reporting both studies simply as “there was no significant difference” makes them sound almost identical. They are not. The first leaves considerable uncertainty about educationally important differences. The second provides substantially tighter information about their magnitude.
06 · What This Means for You
Match Your Conclusion to the Question Your Analysis Actually Answered
If your ordinary significance test is nonsignificant, resist replacing the statistical statement with the stronger substantive conclusion “there is no difference.” First determine how much uncertainty remains.
Ask what differences would matter in the context of your research. Then compare that range with the effects still compatible with your data. This helps distinguish a result suggesting little meaningful effect from one that remains too uncertain to know .
A simple decision framework
If the result is nonsignificant and the uncertainty interval is wide
Report that the study did not establish a difference and describe the remaining uncertainty.
If meaningful differences remain compatible with the data
Do not conclude that the groups are meaningfully similar.
If your research question is explicitly about similarity or negligible effects
Define defensible equivalence bounds and use an analysis designed to evaluate equivalence.
If a precise result excludes differences large enough to matter
State what magnitude has been excluded rather than claiming absolute equality.
The language should reflect the evidence. “We found no statistically significant difference” may be accurate when that is literally the result. “The groups did not differ” is a stronger claim. “The results were compatible only with differences smaller than our prespecified threshold of practical importance” is stronger still, but it can be justified when the analysis genuinely supports it.
07 · A Quick Checklist
Before Writing “There Was No Difference”
Before interpreting a nonsignificant comparison, check:
Report the estimated difference rather than only the p-value.
Examine the uncertainty interval around the estimated effect.
Determine what magnitude of difference would matter substantively.
Check whether differences of that magnitude remain compatible with the data.
Consider whether the study was designed to provide informative estimates for effects of interest.
Assess whether measurement, missing data, bias, or model assumptions weaken the inference.
If similarity is the research question, consider equivalence testing rather than relying on nonsignificance.
Make sure your written conclusion does not claim more than the analysis supports.
09 · The Bottom Line
Nonsignificance Does Not Establish No Difference
The Bottom Line
“No significant difference” does not show that there is no difference when meaningful differences remain compatible with the evidence. A nonsignificant conventional test is a failure to reject the null hypothesis, not a demonstration that the groups are equal.
Interpret the estimated effect together with its uncertainty and a defensible definition of what difference would matter. If the real research question is whether two conditions are sufficiently similar, use evidence and methods capable of answering that question rather than treating p >.05 as proof of sameness.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation