Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Did You Mistake Statistical Significance for Evidence Quality?

A small p-value does not tell you whether an effect is large, important, unbiased, precise enough for a decision, or supported by high-certainty evidence. Statistical significance answers a much narrower question.

853
Statistical Significance vs. Evidence Quality Guide 853 of 899
01 · The Question

Does a Statistically Significant Result Mean the Evidence Is Strong?

A study reports p < 0.05. The result receives an asterisk in the table, the abstract calls the finding significant, and the discussion may treat it as evidence that the hypothesis was supported.

It is easy to take the next step: statistically significant means convincing evidence.

That step is too large.

Statistical significance addresses how the observed data relate to a specified statistical model and null hypothesis. It does not, by itself, tell you whether the estimated effect is large enough to matter, whether the study is biased, whether the evidence applies to your population, whether the result is precise enough for a decision, or whether the wider body of evidence deserves high confidence.

02 · The Short Answer

A P-Value Is Not a Quality Score

In Brief

Statistical significance does not establish evidence quality: interpret the effect estimate and its uncertainty, then separately consider risk of bias, consistency, directness, imprecision, missing evidence, and practical importance.

A small p-value can accompany a trivial effect, a biased study, or evidence of limited relevance. A result above a conventional significance threshold can still be compatible with an important effect when the estimate is imprecise. The threshold does not make these judgments for you.

03 · What You Need to Know

Statistical Significance Answers a Narrower Question Than Researchers Often Assume

What does a p-value actually tell you?

In conventional null-hypothesis significance testing, a p-value describes how incompatible the observed data are with a specified statistical model, including the null hypothesis and its assumptions. It is not the probability that your hypothesis is true. It is also not the probability that the result occurred "by chance."

The American Statistical Association explicitly cautions that p-values do not measure the probability that the studied hypothesis is true and that scientific conclusions should not depend solely on whether a p-value crosses a particular threshold.

The 0.05 threshold is not a border between truth and falsehood

Results just below and just above 0.05 are not fundamentally different forms of evidence. A result with p = 0.049 does not become scientifically compelling merely by crossing the conventional threshold, while p = 0.051 does not suddenly become evidence of no effect.

Cochrane advises review authors not to describe results simply as "statistically significant," "not statistically significant," or "non-significant," and instead recommends focusing on effect estimates, confidence intervals, and exact p-values where these are reported.

Watch Out

"Not statistically significant" does not mean "no effect." If the estimate is imprecise, the data may still be compatible with meaningful benefit, meaningful harm, and little or no effect.

Effect size asks a different question

Suppose a very large study detects a difference of 0.2 points on a 100-point scale with p < 0.001. The result may be highly statistically incompatible with the exact null hypothesis while remaining too small to matter educationally, clinically, socially, or practically.

The American Statistical Association emphasizes that statistical significance does not measure the size or importance of an effect. Cochrane likewise warns that large studies can produce small p-values for effects that are trivial in practical terms.

Statistical significance Conventionally describes whether a statistical test produces a p-value below a selected threshold under the specified model.
Practical importance Concerns whether the magnitude of an effect is large enough to matter for the scientific, clinical, educational, policy, or other decision being considered.

Confidence intervals show information a significance label hides

A point estimate tells you the observed magnitude and direction of an effect. A confidence interval communicates statistical uncertainty around that estimate. Wider intervals generally indicate greater uncertainty, while narrower intervals indicate greater precision.

This can reveal why binary significance labels are inadequate. Two studies may estimate the same effect but differ in precision. One may cross the conventional significance threshold and the other may not, even though their results are quite compatible. Cochrane therefore recommends interpreting the estimate together with its confidence interval.

Question Can statistical significance answer it? What else is needed?
Is the estimated effect large? No Effect size in meaningful units
Is the effect practically important? No Context and a meaningful decision threshold or interpretation
Is the estimate precise? Not adequately by itself Confidence or credible interval and relevant decision thresholds
Is the study unbiased? No Design-specific risk-of-bias assessment
Does the evidence apply to my population? No Assessment of directness and context
Is the body of evidence consistent? No Comparison and synthesis across relevant studies
Is the overall evidence high certainty? No A structured certainty assessment considering multiple domains

A small p-value cannot repair a biased study

Statistical tests operate on the data produced by the study. If those data are systematically distorted by selection bias, confounding, missing data, measurement problems, selective reporting, or another source of bias, an extremely small p-value does not make the underlying design more credible.

Cochrane treats risk of bias as a separate methodological consideration precisely because statistical precision and systematic error are different problems. A highly precise estimate can still be biased.

Sample size can change statistical significance without changing the effect itself

Statistical significance depends partly on precision, and precision is often strongly influenced by sample size. As sample size increases, even small departures from the null can become statistically detectable.

Conversely, a small study can estimate a substantial effect but fail to cross a conventional significance threshold because the estimate is too imprecise. Cochrane specifically notes that p-values depend on both the effect estimate and its precision.

Statistical significance is not the same as certainty of evidence

Certainty concerns how confident you should be in a body of evidence for a particular outcome. In the GRADE framework used by Cochrane, certainty is evaluated using considerations including risk of bias, consistency of effects, imprecision, indirectness, and publication bias.

A statistically significant meta-analysis can therefore still provide low-certainty evidence if, for example, the included studies have serious methodological limitations or the evidence is indirect. Conversely, evidence can be informative even when a conventional significance threshold is not crossed, particularly when the confidence interval is sufficiently precise to address effects that matter.

Statistical significance does not determine which evidence deserves more weight

Giving "significant" studies more rhetorical weight than "non-significant" studies creates a distorted synthesis. The studies should instead be compared according to their estimates, uncertainty, methods, relevance, and contribution to the complete evidence base.

This is why stronger and weaker evidence should be distinguished using substantive methodological criteria rather than a significance threshold.

Be careful when studies disagree only because their significance labels differ

One study may report p = 0.03 and another p = 0.08. It is tempting to say the first found an effect and the second contradicted it. That conclusion does not follow from the two thresholds alone.

The effect estimates and intervals may be similar. If so, the apparent disagreement is partly an artifact of dichotomizing continuous statistical evidence. Before describing results as contradictory, compare what the studies actually estimated.

04 · A Practical Example

When the Smaller P-Value Gives the Less Important Result

Hypothetical Example

Two studies of a digital learning intervention

Two studies measure examination performance on a 100-point scale. Study A is very large and reports an average improvement of 0.4 points with a very small p-value. Study B is much smaller and estimates an improvement of 4 points, but its confidence interval is wide and includes effects ranging from negligible to educationally important.

Ignore the significance contest The researcher does not conclude that Study A provides the more important finding merely because its p-value is smaller.
Examine magnitude Study A estimates a very small effect. Whether 0.4 points matters requires substantive interpretation.
Examine precision Study B estimates a larger effect but with substantial uncertainty, so the true effect compatible with the data spans a wider range.
Appraise bias and relevance Both studies are evaluated for methodological limitations and applicability independently of their p-values.
Interpret the evidence The researcher reports what each estimate suggests and how uncertain it is rather than declaring one study "positive" and the other "negative."

Neither result can be understood adequately from its significance label alone. One is precise but potentially trivial. The other may be important but uncertain. The interesting part begins where the asterisks stop.

05 · What Researchers Often Get Wrong

Common Misinterpretations of Statistical Significance

Misconception

Does P < 0.05 Mean There Is a 95% Chance the Hypothesis Is True?

No. A p-value does not provide the probability that the research hypothesis is true or false. The American Statistical Association identifies this as a fundamental misinterpretation.

Misconception

Does Statistical Significance Mean the Effect Is Important?

No. A sufficiently precise study can produce a small p-value for an effect that is too small to matter in practice. Effect magnitude and substantive importance must be evaluated separately.

Misconception

Does P > 0.05 Mean There Is No Effect?

No. A result may be too imprecise to distinguish among no effect, meaningful benefit, and meaningful harm. Cochrane specifically warns against confusing lack of evidence of an effect with evidence of no effect.

Misconception

Is P = 0.049 Meaningfully Stronger Than P = 0.051?

Not merely because the two values fall on opposite sides of 0.05. Treating the threshold as a sharp scientific boundary discards information and can create artificial disagreement between otherwise similar results.

Misconception

Does a Smaller P-Value Mean a Larger Effect?

No. The p-value reflects both effect magnitude and precision. A tiny effect estimated from a very large sample can produce a smaller p-value than a much larger but less precisely estimated effect.

Misconception

Does a Significant Meta-analysis Mean High-Certainty Evidence?

No. Certainty also depends on concerns such as risk of bias, inconsistency, indirectness, imprecision, and publication bias. A pooled result can cross a significance threshold while the underlying body of evidence still warrants limited confidence.

06 · What This Means for You

Move From "Is It Significant?" to "What Does the Evidence Actually Support?"

When reading a statistical result, resist beginning and ending with the p-value. Start with the estimate itself, examine its uncertainty, and then place that result inside the methodological and substantive context of the study.

A simple decision framework

If p < 0.05
Examine the effect magnitude, confidence interval, methodological credibility, and practical importance before deciding what the finding means.
If p > 0.05
Examine whether the interval rules out effects that would matter before concluding that there is little or no meaningful effect.
If the effect is statistically detectable but very small
Separate statistical evidence against the null from the question of whether the magnitude matters.
If the effect appears large but is imprecise
Communicate the uncertainty rather than treating the point estimate as established.
If the study has serious methodological limitations
Do not allow a small p-value to override the risk-of-bias assessment.
If several studies differ in significance labels
Compare their effect estimates and uncertainty before concluding that their findings conflict.

Statistical results also need proportionate language. When uncertainty or methodological limitations remain substantial, avoid making the evidence sound more certain than it really is.

07 · A Quick Checklist

Before You Call a Statistically Significant Finding Strong Evidence

For each important statistical result, check:
I examined the effect estimate rather than relying only on whether the p-value crossed 0.05.
I examined the confidence interval or another appropriate representation of uncertainty.
I distinguished statistical significance from practical, clinical, educational, or substantive importance.
I did not interpret p > 0.05 automatically as evidence of no effect.
I did not interpret a smaller p-value automatically as a larger or more important effect.
I considered risk of bias independently of the statistical significance of the result.
I considered whether the evidence is direct and consistent with other relevant studies.
I kept statistical significance separate from the overall certainty of the evidence.
08 · Frequently Asked Questions

Questions About P-Values, Significance, and Evidence Quality

What does p < 0.05 actually mean?

Under the statistical model and null hypothesis being tested, it means the calculated p-value is below 0.05. It does not mean there is a 95% probability that the research hypothesis is true, nor does it establish that the effect is important.

Is a statistically significant result strong evidence?

It may contribute evidence against a specified null hypothesis, but evidence strength requires additional judgments. Risk of bias, effect magnitude, precision, directness, consistency, and missing evidence can all materially affect how much confidence the result deserves.

What is the difference between statistical and practical significance?

Statistical significance concerns a statistical test relative to a threshold. Practical significance concerns whether the estimated magnitude is consequential in the real context of the research question. One does not establish the other.

Does a non-significant result prove the null hypothesis?

No. A study may simply lack sufficient precision to distinguish no effect from effects that would matter. Examine the estimate and its uncertainty before deciding what the evidence excludes or remains compatible with.

Why should I report confidence intervals?

They show the uncertainty surrounding an estimate and help readers judge which effect sizes remain compatible with the data. This provides substantially more information than a binary significant or non-significant label.

Can a tiny effect be statistically significant?

Yes. Large samples can estimate small effects precisely enough to produce very small p-values. Whether that effect matters requires substantive interpretation rather than a significance threshold.

Can an important effect be statistically non-significant?

Yes. A point estimate may suggest an important effect while the available data remain too imprecise to distinguish it confidently from smaller effects or no effect. The appropriate conclusion is uncertainty, not automatically "no effect."

Is statistical significance part of GRADE certainty?

Not as a standalone quality criterion. GRADE considers imprecision alongside risk of bias, inconsistency, indirectness, and publication bias. Whether a p-value crosses 0.05 does not by itself determine the certainty rating.

09 · The Bottom Line

A Small P-Value Cannot Tell You How Much the Evidence Deserves Your Confidence

The Bottom Line

Statistical significance is not evidence quality: interpret the magnitude and uncertainty of the effect, then judge the credibility and certainty of the evidence using the methodological information that a p-value cannot provide.

A threshold can summarize one feature of a statistical test, but it cannot tell you whether a finding is important, unbiased, applicable, or certain. Those judgments still belong to the researcher.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes