01 · The Question
What If the Literature Keeps Asking Whether an Effect Exists Instead of Whether It Matters?
You review a mature literature and encounter the same language repeatedly: significant, non-significant, significant, significant again.
At first, this can make the evidence appear decisive. If enough studies report p <.05, surely the important question has been answered.
Not necessarily.
A significance test can contribute information about how compatible observed data are with a specified statistical model or null hypothesis. It does not, by itself, tell you whether an effect is large, useful, consequential, precisely estimated, theoretically important, or worth acting on. The American Statistical Association has explicitly cautioned against treating a p-value or a statistical-significance threshold as a measure of the size or importance of an effect.
A literature organized primarily around whether effects cross a significance threshold can therefore accumulate publications while leaving a more consequential question unresolved: how much difference does the phenomenon actually make?
03 · What You Need to Know
Statistical Detectability and Substantive Importance Are Different Questions
A p-value does not tell you how large an effect is
The American Statistical Association's statement on p-values emphasizes that statistical significance does not measure the size of an effect or the importance of a result. This distinction is foundational.
Suppose an educational intervention improves examination performance by 0.2 percentage points in a very large sample and the result produces a very small p-value. The data may provide evidence against a precise null hypothesis of no difference. That does not tell you whether a 0.2-point improvement matters educationally.
Conversely, a smaller study could estimate an improvement large enough to matter but with substantial uncertainty, producing a p-value above.05. Labeling the first result a “success” and the second “no effect” would obscure what the studies actually estimated.
Statistical significance
A threshold-based interpretation of a statistical test, commonly based on whether a p-value falls below a conventional cutoff.
Meaningful effect
An effect large enough to matter for the substantive scientific, practical, clinical, educational, policy, or other decision under consideration.
Sample size can change significance without changing the substantive effect
P-values depend partly on precision, and precision is often strongly influenced by sample size. As samples grow, smaller effects can become statistically detectable.
This is not a defect in statistics. It becomes a problem when researchers treat detectability as importance.
Cochrane's guidance on interpreting evidence makes this point directly: a small p-value can accompany an effect too trivial to produce an important benefit. Cochrane therefore advises review authors to focus interpretation on effect estimates and confidence intervals rather than relying on a binary distinction between statistically significant and non-significant findings.
The reverse problem occurs with small samples. A potentially consequential effect may fail to cross a significance threshold simply because the estimate is imprecise.
Effect estimates answer the question significance tests cannot
If your substantive question is “How much does X change Y?”, the estimated magnitude of the effect needs to be central to the analysis.
Depending on the research question and outcome, researchers might report mean differences, standardized mean differences, risk differences, risk ratios, odds ratios, correlations, regression coefficients, rate ratios, or other appropriate effect measures.
The specific metric matters less than the underlying principle: report the quantity that represents the effect you want to understand and interpret its magnitude in context.
An effect estimate without context can still be difficult to interpret. A standardized effect of 0.30, for example, does not become “small” in every field merely because a conventional label says so. Whether an effect matters depends on the outcome, baseline conditions, costs, risks, feasibility, available alternatives, and what differences are consequential in that domain.
Confidence intervals help reveal what remains plausible
A point estimate is not known with perfect precision. Confidence intervals provide information about sampling uncertainty around that estimate under the statistical procedure used.
Cochrane recommends interpreting point estimates together with confidence intervals because the interval can reveal whether the data remain compatible with importantly different conclusions.
Suppose a study estimates an improvement of four points, with a confidence interval extending from one to seven points. If a three-point improvement would be substantively important, the interval includes both effects below and above that threshold. The study therefore leaves uncertainty about whether the true effect is large enough to matter.
This is more informative than simply observing whether the interval excludes zero.
| What the literature reports |
What it establishes |
What may remain unanswered |
| p <.05 |
The observed data meet a chosen significance criterion under the specified analysis |
How large and consequential the effect is |
| p >.05 |
The chosen significance criterion was not met |
Whether important effects remain plausible because of imprecision |
| Effect estimate without uncertainty |
An estimated magnitude |
How precisely that magnitude is known |
| Effect estimate and confidence interval |
Magnitude plus sampling uncertainty |
Whether the plausible effects are meaningful in the relevant context |
| Very precise tiny effect |
A small effect estimated with considerable precision |
Whether the difference matters enough to influence understanding or action |
“Non-significant” does not mean “no effect”
A study that fails to cross a conventional significance threshold has not necessarily demonstrated the absence of an effect.
Cochrane specifically warns against confusing lack of evidence of an effect with evidence of a lack of effect. A wide confidence interval might include no effect, a small effect, and a substantial effect simultaneously.
The appropriate interpretation depends on the estimate and its uncertainty.
If researchers genuinely want to establish that effects are sufficiently small to be considered unimportant, approaches designed around equivalence, non-inferiority, or predefined ranges of practical equivalence may sometimes be more suitable than conventional null-hypothesis significance testing. Which approach is appropriate depends on the substantive question.
A meaningful difference must be justified, not invented after seeing the results
Once researchers recognize that statistical significance is insufficient, another temptation appears: declaring whatever observed effect they obtained “practically significant.”
That solves nothing.
A threshold for meaningful change should ideally be justified independently of the observed result. Depending on the field, evidence may come from established minimal important differences, stakeholder judgments, validated benchmarks, policy thresholds, cost-benefit considerations, prior empirical work, or theoretically defensible criteria.
In some areas, no accepted threshold exists. Researchers should say so rather than manufacture one.
Watch Out
Replacing an arbitrary p <.05 threshold with an equally arbitrary “meaningful effect” threshold does not solve the interpretive problem. The substantive criterion itself needs justification.
Meaningfulness may differ across stakeholders and contexts
A modest effect can be important when an intervention is inexpensive, safe, scalable, and applied to millions of people. The same effect may be unattractive when achieving it requires substantial cost, risk, workload, or displacement of a better alternative.
Likewise, what counts as meaningful to an individual participant may differ from what matters to an institution, policymaker, clinician, teacher, or researcher testing a theory.
This is why practical significance should not be reduced to a universal table of effect-size labels.
The outcome itself must also matter
A large effect on a trivial or indirect outcome does not automatically become consequential.
If studies detect substantial changes in engagement but the real decision concerns durable learning, the unresolved issue may be that researchers have measured an outcome that does not adequately answer the substantive question.
Magnitude and outcome relevance therefore need to be considered together.
Statistical significance can hide the real research gap
Imagine 40 studies demonstrating statistically significant effects of an intervention. A conventional review might conclude that another study is unnecessary.
But suppose most papers report little about magnitude, confidence intervals are rarely interpreted substantively, and nobody has established whether the observed differences exceed a threshold that matters in practice.
The unanswered question is no longer simply whether the intervention has an effect. It is whether the effect is consequential enough to justify the conclusions or decisions being made.
This is a good example of why research uncertainty can persist even when studies themselves are not missing.