01 · The Question
If a Result Is Statistically Significant, Does It Actually Matter?
A paper reports P < 0.001. The authors emphasize the finding, perhaps describing a clear difference or association. It is tempting to assume that such a small P-value means something important has been discovered.
But statistical significance and substantive importance answer different questions. A result can provide strong evidence against a particular null hypothesis while the estimated effect is so small that it has little practical, clinical, educational, theoretical, or policy relevance.
Critical appraisal therefore requires another question after you see a statistically significant result: how large is the effect, how precisely was it estimated, and would an effect of that magnitude actually matter in this context?
03 · What You Need to Know
Statistical Significance Is Not a Measure of Importance
What statistical significance actually tells you
In conventional null-hypothesis significance testing, a P-value summarizes how incompatible the observed data are with a specified statistical model, usually including a null hypothesis. If the P-value falls below a prespecified significance level such as 0.05, researchers traditionally describe the result as statistically significant.
That designation does not measure effect magnitude. The American Statistical Association explicitly cautions that a P-value or statistical significance does not measure the size of an effect or the importance of a result. It also advises against basing scientific conclusions solely on whether a P-value crosses a particular threshold.
This distinction is fundamental. Evidence that an effect differs detectably from an exact null value is not the same as evidence that the difference is large enough to matter.
Statistical significance
Describes a result relative to a statistical testing procedure and specified null model.
Substantive importance
Asks whether the magnitude and consequences of an effect matter for the scientific, clinical, educational, practical, or other question being studied.
Why sample size changes the picture
The P-value reflects not only the estimated effect but also its statistical precision. Precision generally increases as the amount of relevant information increases. Consequently, a very large study may detect an effect that is extremely small but estimated very precisely.
Cochrane makes this point explicitly: in a large study, a small P-value can represent detection of a trivial effect. The point estimate and confidence interval are therefore essential to interpretation.
This is why very large samples can make trivial effects look statistically impressive. The statistical procedure may be doing exactly what it is supposed to do. The mistake occurs when the reader interprets detectability as importance.
The effect estimate tells you what happened on the relevant scale
Suppose a study of an educational intervention reports P < 0.001 for the difference in examination performance. That sounds impressive until you learn that the estimated improvement is 0.3 points on a 100-point scale.
The P-value does not answer whether 0.3 points matters. You need the effect estimate for that question, ideally expressed in units that make substantive interpretation possible.
Depending on the study, useful measures may include mean differences, risk differences, risk ratios, odds ratios, correlations, regression coefficients, standardized mean differences, or other estimands. The key is to determine what the reported effect represents in the actual context.
Absolute and relative effects can tell different stories
Relative effects can appear impressive while corresponding to modest absolute changes. Suppose an intervention reduces the probability of an outcome from 2 in 10,000 to 1 in 10,000. That is a 50% relative reduction but an absolute reduction of only 1 event per 10,000 people.
Whether that is worthwhile cannot be determined from either percentage alone. The answer could depend on the severity of the outcome, cost and burden of the intervention, adverse effects, feasibility, and the number of people affected.
For decision-making, both relative and absolute measures may therefore be useful. An effect that sounds large on one scale can look quite different on another.
The confidence interval tells you how precisely importance has been estimated
A point estimate should not be interpreted without its uncertainty. A confidence interval shows the range of values supported by the data and statistical model at the stated confidence level and helps reveal how precisely the effect has been estimated.
Suppose an intervention produces an estimated improvement of 1 point with a 95% confidence interval from 0.9 to 1.1. If an improvement would need to reach 5 points before it becomes substantively worthwhile, the interval is precise but centered on an effect too small to meet that criterion.
Compare that with an estimated improvement of 1 point and a confidence interval from -3 to 7. The point estimate is still small, but the study is too imprecise to exclude effects that might matter. Those two studies should not receive the same interpretation.
This is why P-values, effect estimates, and confidence intervals should be interpreted together rather than allowing the significance label to dominate the result.
Importance requires a meaningful reference point
Calling an effect important requires some basis for deciding what magnitude matters. That basis should ideally come from the research context rather than from the observed P-value.
In clinical research, researchers may consider a minimal clinically important difference. In education, a difference might be interpreted against meaningful changes in achievement, progression, workload, or implementation cost. Other fields may use theoretical predictions, stakeholder judgments, economic consequences, policy thresholds, or other domain-specific criteria.
The relevant threshold does not always have to be a single universally accepted number. Often there is legitimate uncertainty or disagreement about what constitutes a meaningful effect. Critical appraisal should make that disagreement visible rather than pretending statistical significance resolves it.
A small effect can still matter
The opposite shortcut is also dangerous. "Small" does not automatically mean "unimportant."
A modest individual-level effect could matter substantially when an intervention is inexpensive, low risk, scalable, and applied across a very large population. Small changes in common outcomes can accumulate into meaningful population-level consequences. Conversely, a larger effect may not justify an intervention that is expensive, burdensome, risky, or difficult to implement.
Importance is therefore relational. It depends on the effect, outcome, population, alternatives, costs, benefits, harms, and decision being considered.
Standardized effect-size labels are not universal importance thresholds
Researchers sometimes classify standardized effects using generic labels such as small, medium, or large. These conventions can provide rough orientation, but they should not be treated as universal definitions of importance.
An effect that is numerically small according to a generic benchmark may be consequential in one field and negligible in another. Whenever possible, interpret the effect using domain knowledge and meaningful units rather than outsourcing substantive judgment to a conventional label.
Statistical significance cannot repair bias
Even a large and statistically significant effect may be untrustworthy if the study is seriously biased. Confounding, attrition, poor measurement, selective reporting, inappropriate statistical adjustment, or analytical flexibility can all undermine interpretation.
A small P-value describes the statistical result under the analysis that produced it. It does not certify the research design.
This becomes especially important when a paper contains many statistical tests or there is reason to suspect selective reporting of significant analyses. A nominally significant result may look much less persuasive once the broader analytical process is considered.
Watch Out
Do not replace one mechanical rule with another. "Statistically significant means important" is wrong, but so is "small effect means unimportant." Importance requires a substantive judgment about magnitude and consequences in the context of the research question.
04 · A Practical Example
When P < 0.001 Is Less Impressive Than It Looks
Hypothetical Example
A learning platform produces a statistically detectable improvement
Suppose a very large randomized study evaluates a new learning platform. Students using the platform score an average of 80.4 on a standardized assessment, compared with 80.0 in the control group. The estimated difference is 0.4 points on a 100-point scale, with a 95% confidence interval from 0.25 to 0.55 and P < 0.001.
Read the statistical evidence. The very small P-value indicates that the observed difference is difficult to reconcile with an exact zero-effect null under the assumptions of the statistical test.
Read the magnitude. The estimated improvement is 0.4 points on a 100-point assessment. Statistical detectability does not establish that this change is educationally consequential.
Read the uncertainty. The confidence interval is narrow, from 0.25 to 0.55. The study has therefore estimated this small effect relatively precisely rather than merely producing an uncertain estimate centered near zero.
Apply a substantive criterion. Suppose independent educational considerations suggest that improvements smaller than 2 points would not justify the financial and instructional costs of adopting the platform. The entire confidence interval remains below that benchmark.
Interpret the finding. The evidence supports the existence of a small positive effect under the study's assumptions, but it does not support the stronger claim that the effect is large enough to justify implementation under the stated criterion.
Notice that nothing is statistically contradictory here. The result can simultaneously be highly statistically significant, precisely estimated, and too small to matter for a particular decision.