01 · The Question
Which Statistical Result Deserves Your Attention?
A paper reports a treatment effect of 4.2 points, a 95% confidence interval of 0.3 to 8.1, and P = 0.034. Which number tells you whether the finding matters?
None of them answers that question completely by itself. The effect estimate tells you how large the observed difference is. The confidence interval tells you about uncertainty around that estimate. The P-value summarizes how incompatible the observed data are with a specified statistical model, typically one involving a null hypothesis.
The mistake is treating these quantities as competitors and choosing one number to carry the entire interpretation. They answer related but distinct questions.
03 · What You Need to Know
Each Quantity Answers a Different Statistical Question
The effect estimate asks: how large is the observed effect?
An effect estimate describes the magnitude and direction of an observed relationship, difference, or association. Depending on the research question, this might be a mean difference, risk difference, risk ratio, odds ratio, hazard ratio, correlation coefficient, regression coefficient, standardized mean difference, or another estimand.
Suppose an intervention reduces a mean outcome by 0.2 points. Knowing that P < 0.001 does not tell you whether a reduction of 0.2 points matters. That requires understanding the effect estimate in its substantive context.
Effect magnitude may be expressed on an absolute scale or a relative scale. The choice can materially affect interpretation. A relative reduction of 50%, for example, could represent a decrease in risk from 20% to 10% or from 0.02% to 0.01%. The relative change is the same, but the absolute consequences are very different.
The confidence interval asks: how uncertain is the estimate?
A point estimate is only one estimate obtained from a sample. A confidence interval expresses statistical uncertainty around that estimate under the assumptions of the procedure used to construct it.
For a conventional 95% frequentist confidence interval, the formal interpretation concerns the long-run performance of the interval-generating procedure: if the procedure were repeated across many comparable samples, 95% of the resulting intervals would contain the true parameter under the model assumptions. It is therefore imprecise to say that there is a 95% probability that the fixed true parameter lies inside one particular frequentist interval.
In practical appraisal, interval width is especially informative. A narrow interval suggests greater precision, while a wide interval may leave several substantively different effects compatible with the data. Cochrane explicitly recommends considering point estimates together with their confidence intervals when interpreting intervention effects.
The P-value asks a narrower question
A P-value is calculated relative to a statistical model and hypothesis. In a conventional null-hypothesis significance test, it represents the probability, assuming the null hypothesis and other model assumptions used in the calculation, of obtaining a result at least as incompatible with that hypothesis as the observed result.
That definition rules out several common interpretations. A P-value of 0.03 does not mean there is a 3% probability that the null hypothesis is true. It does not mean there is a 97% probability that the research hypothesis is correct. It does not tell you the probability that the result occurred "by chance." And it does not measure the size or practical importance of the effect.
Quantity
Main question it helps answer
What it does not tell you by itself
Effect estimate
How large and in what direction is the observed effect?
How precise the estimate is
Confidence interval
What range of parameter values is compatible with the data and model at the stated confidence level?
Whether every value in that range is substantively important
P-value
How incompatible are the observed data with the specified null model, according to the chosen test?
The size, importance, or probability that the hypothesis is true
Why P < 0.05 should not become the conclusion
The familiar 0.05 threshold is a convention, not a boundary separating real effects from unreal ones. Results with P = 0.049 and P = 0.051 provide nearly the same numerical evidence, yet dichotomizing them can make their interpretations sound dramatically different.
Cochrane guidance consequently advises against undue reliance on labels such as statistically significant and nonsignificant and recommends focusing interpretation on estimates of effect and confidence intervals, with exact P-values reported when used.
The American Statistical Association has likewise emphasized that scientific conclusions should not be based only on whether a P-value crosses a particular threshold and that statistical significance does not measure the size or importance of an effect.
A small P-value can accompany a tiny effect
P-values depend partly on precision, which is strongly influenced by sample size. With sufficiently large samples, very small effects can generate very small P-values.
Imagine a digital intervention that improves an outcome by 0.05 points on a 100-point scale. In an enormous dataset, the estimate might be extremely precise and yield P < 0.001. The evidence may be strong that the population effect is not exactly zero, while the magnitude remains practically negligible.
This is why a statistically significant result can still be unimportant and why very large samples can make trivial effects appear statistically impressive .
A large P-value does not demonstrate that there is no effect
The reverse mistake is equally common. P > 0.05 does not establish that the null hypothesis is true or that the groups are equivalent. A study may simply be too imprecise to distinguish among several plausible effects.
Suppose an estimated mean difference is 5 points with a 95% confidence interval from -2 to 12. The interval includes zero, but it also includes potentially meaningful positive effects. Describing this only as "no significant difference" discards important information about what the study can and cannot exclude.
This becomes especially important when evaluating whether a nonsignificant result is nevertheless informative or whether the study may be too imprecise to answer its main question convincingly .
Confidence intervals connect magnitude and uncertainty
One reason confidence intervals are so useful is that they force you to consider more than whether the null value is included. Ask what the entire interval permits.
For a mean difference, the null value is typically 0. For ratio measures such as a risk ratio or odds ratio, it is typically 1. Whether the interval crosses that null value relates to conventional statistical significance when the confidence interval and hypothesis test use corresponding methods. But stopping there wastes much of the information the interval contains.
Consider an estimated risk ratio of 0.80 with a 95% confidence interval from 0.77 to 0.83. That tells a different evidential story from a risk ratio of 0.80 with an interval from 0.45 to 1.42, even though the point estimates are identical. The second study leaves much greater uncertainty about the underlying effect.
Effect size does not automatically mean practical importance
The phrase "effect size" can create another shortcut. A standardized effect of a certain numerical magnitude is not automatically small, medium, or large in every discipline or application. Context determines whether an effect matters.
A seemingly small effect may be important when an intervention is inexpensive, scalable, low risk, and applied to millions of people. A numerically larger effect may matter little if it occurs on an outcome with minimal practical relevance or comes with substantial costs and harms.
Whenever possible, interpret effects using domain knowledge, meaningful units, baseline risk, clinically or practically important thresholds, costs, harms, and consequences rather than relying mechanically on generic magnitude labels.
All three quantities inherit the weaknesses of the underlying analysis
A beautifully narrow confidence interval around a biased estimate is still a precisely estimated biased result. A tiny P-value from an inappropriate model does not rescue the model. A large effect estimate generated by selective reporting remains vulnerable to selective reporting.
Statistical summaries should therefore be interpreted only after asking whether the design, measurement, analysis, and reporting are credible. This is also why successfully recalculating the reported statistics cannot by itself establish that the study's inference is valid.
Watch Out
Do not read a confidence interval merely as a more elaborate significance test. Asking only whether it crosses the null value turns a range of information about magnitude and precision back into the same binary P < 0.05 decision you were trying to move beyond.
04 · A Practical Example
How the Same Result Changes When You Read All Three Numbers
Hypothetical Example
An educational intervention improves examination scores
Suppose a randomized study compares a new learning intervention with usual instruction. The adjusted mean difference in the final examination score is 2.0 points on a 100-point scale, with a 95% confidence interval from 0.4 to 3.6 points and P = 0.015.
Read the P-value. P = 0.015 indicates that the observed result is relatively incompatible with the specified null model under the assumptions of the test. It does not tell you whether a 2-point improvement matters educationally.
Read the effect estimate. The estimated difference is 2 points. Now the substantive question becomes unavoidable: is a 2-point improvement meaningful on this assessment?
Read the confidence interval. The interval from 0.4 to 3.6 points shows the uncertainty around the estimate. Effects near the lower end may be negligible, while effects near the upper end might be more consequential, depending on the educational context.
Add substantive knowledge. Suppose researchers had good reason, established independently of the observed result, to regard a 5-point difference as the smallest effect that would justify the cost of implementing the intervention. The entire reported interval lies below that benchmark.
Interpret rather than label. The study provides evidence against an exact zero-effect null under its statistical model, but the estimated improvement and its confidence interval would not support the claim that the intervention achieves the stipulated 5-point threshold.
The example shows why "statistically significant" is not a sufficient interpretation. The P-value answers one inferential question, while the effect estimate, confidence interval, and substantive threshold address questions that are usually closer to the decision a researcher actually cares about.
06 · What This Means for You
Read the Estimate First, Then Its Uncertainty, Then the P-Value in Context
When critically appraising a conventional statistical result, a useful reading sequence is to begin with the estimated effect and ask what it means in the real units or substantive context of the study. Then examine the confidence interval to understand how precisely that effect has been estimated. Finally, use the P-value, when reported, as additional information about the specified statistical test rather than as the verdict on the study.
A simple interpretation framework
If the effect estimate is substantively important and the interval is reasonably precise
Ask whether the design and analysis are credible enough for that estimate to support the claimed inference.
If P is small but the estimated effect is trivial
Do not confuse strong evidence against an exact null with evidence of practical importance.
If P is large and the confidence interval is wide
Treat the result as imprecise rather than concluding automatically that there is no effect.
If the confidence interval excludes effects large enough to matter
That may be substantively informative even if the usual significance label does not capture the point you care about.
The hierarchy is not absolute. Some research questions use different inferential frameworks, and not every study reports all three quantities. The general principle is to extract as much information as the analysis legitimately provides without allowing one threshold to replace scientific interpretation.
07 · A Quick Checklist
How to Read a Statistical Result Without Stopping at P < 0.05
When interpreting a statistical result, check:
What effect or parameter is actually being estimated?
What is the magnitude and direction of the effect in substantively meaningful terms?
Is an absolute effect available as well as a relative or standardized effect where that would improve interpretation?
How wide is the confidence interval, and what substantively different effects does it include?
What null hypothesis and statistical model generated the P-value?
Are you interpreting the P-value as evidence about compatibility with the null model rather than as the probability that a hypothesis is true?
Have you separated statistical significance from practical, clinical, educational, or theoretical importance?
Are the design and analysis credible enough for any of these statistical summaries to support the claimed inference?
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation