01 · The Question
Can You Tell From a Published Paper That a Study Was Underpowered?
A paper reports no statistically significant difference. The sample contains 60 participants, the confidence interval is wide, and the authors conclude that the intervention had no effect. Was the study simply too small to detect one?
Possibly, but calling a study underpowered requires more than noticing a small sample or P > 0.05. Statistical power is defined relative to a particular effect size, significance level, statistical test, variability or event rate, and design. A study can have low power to detect a small effect while having ample power to detect a large one.
From a published paper, you therefore need to reconstruct the question the study was designed to answer and then examine whether its achieved information and precision were adequate for effects that would actually matter.
03 · What You Need to Know
Power Is About Detecting a Particular Effect, Not About Sample Size Alone
What does underpowered actually mean?
Statistical power is the probability that a statistical procedure will reject a specified null hypothesis when a particular alternative is true, given the assumptions used in the calculation. In study planning, researchers commonly choose a target power such as 80% or 90% and calculate how much information they need to detect an effect of a specified magnitude at a chosen significance level.
That definition immediately reveals why "this study has only 100 participants" is not enough to diagnose underpowering. The required sample depends on what difference researchers want to detect, the expected variability or event frequency, the statistical method, allocation, clustering or repeated measurement where relevant, and other design features.
CONSORT 2025 recommends reporting the assumptions supporting a randomized trial's sample-size calculation, including the primary outcome, target difference, relevant outcome values or variability, statistical test, significance level, statistical power, and resulting target sample size. These details let readers understand what the study was designed to detect.
Start with the original sample-size calculation
Find the section explaining how sample size was determined. For a well-reported study, you should be able to identify the primary outcome and the effect magnitude used for planning.
Suppose researchers calculated that 300 participants were required to provide 90% power to detect a 5-point difference. That does not mean the study had 90% power to detect every possible effect. It refers to the assumptions and target effect built into that calculation.
Now ask where the target difference came from. Was 5 points supported by prior evidence, a clinically or practically meaningful threshold, or another defensible rationale? A technically correct calculation built around an implausibly large expected effect can produce a deceptively small required sample.
| What to find in the paper |
Why it matters |
| Primary outcome used for planning |
Power is normally tied to a particular outcome and analysis |
| Target effect |
Shows the magnitude the study was designed to detect |
| Expected variability or event rate |
These assumptions influence the required information |
| Significance level |
Affects the rejection threshold and sample-size requirement |
| Target power |
Shows the planned probability of detecting the specified effect under the assumptions |
| Planned sample size or event count |
Provides a benchmark against which achieved information can be compared |
| Allowance for attrition or missing data |
Loss of usable observations can reduce information available for analysis |
Check whether the study actually achieved what it planned
A study may have been adequately planned and still finish with less information than expected. Recruitment can stop early. Participants may withdraw. Outcome data may be missing. Some observations may be excluded from the primary analysis. In event-driven studies, fewer relevant events may occur than anticipated.
Compare the target sample with the number actually analyzed for the primary outcome, not merely the number initially recruited.
If researchers planned for 500 analyzable participants but the primary analysis contains 310, the original power calculation no longer describes the achieved study straightforwardly. The consequence depends on why observations were lost, how much information remains, and the analytical method.
For some designs, the number of events matters more than the number of participants
Raw participant count can be misleading. In survival analysis, for example, precision may depend heavily on the number of observed events. For binary outcomes, rare events can leave an apparently large study with relatively little information about an effect.
A cohort of 20,000 people sounds enormous, but if the outcome occurs only a handful of times, some effect estimates may remain highly uncertain.
Cochrane similarly notes that precision depends not only on sample size. For continuous outcomes it also depends on outcome variability, while for dichotomous and time-to-event outcomes the number or frequency of events is important.
Confidence intervals show what the completed study actually estimated precisely
Once results exist, the confidence interval is one of the most useful indicators of whether the study has narrowed uncertainty enough to answer its question.
Suppose a treatment has an estimated risk ratio of 0.90 with a 95% confidence interval from 0.45 to 1.80. The interval is compatible with a substantial reduction in risk, little effect, and substantial increase in risk. Regardless of the significance label, that result leaves considerable uncertainty.
Contrast that with a risk ratio of 0.99 with an interval from 0.96 to 1.02. If effects within that range would all be too small to matter, the study has answered a useful question even if P > 0.05.
This is why a nonsignificant result can still be informative. The important issue is not simply whether the test rejected an exact null, but what effects the completed study can reasonably distinguish.
A nonsignificant result is not proof that the study was underpowered
Researchers sometimes reason backward: the result was nonsignificant, therefore the study lacked power. That inference does not follow.
An adequately informative study can produce a nonsignificant result because the true effect is small or absent. Conversely, a statistically significant result does not prove that the study was well powered at the planning stage.
Statistical significance is an observed result. Power is a property of a testing procedure under specified alternatives and assumptions. They should not be treated as interchangeable labels.
Be cautious with observed or post hoc power
After a study produces a nonsignificant result, researchers sometimes calculate power again using the observed effect estimate. This is often called observed or post hoc power.
Using the observed effect to calculate power generally provides little useful information about the completed study. CONSORT has long advised that there is little merit in this approach and recommends using confidence intervals to characterize the uncertainty in the observed result. Methodological critiques likewise note the close dependence between observed power and the observed test result.
This does not make every retrospective power calculation meaningless. Researchers may legitimately ask how much power a completed design would have had for an independently specified effect of substantive interest. That is a different question from inserting the observed effect into a power formula and using the resulting number to explain the P-value.
Watch Out
Do not diagnose an underpowered study by calculating power from its observed effect and P-value. For critical appraisal after the data have been collected, examine the original design assumptions and the precision of the resulting estimates.
Large attrition can undermine more than statistical power
If a study planned to analyze 400 participants but obtains primary-outcome data for only 250, reduced precision is one concern. Attrition can also introduce bias if missingness differs between groups or relates to outcomes.
That distinction matters. Increasing the sample size can address lack of information, but it does not automatically repair systematic bias. A very precise biased estimate remains biased.
Multiple primary outcomes and subgroup analyses complicate power
A sample-size calculation usually targets a particular primary analysis. It does not imply that every secondary outcome, interaction, or subgroup analysis is equally well powered.
Subgroup analyses can be particularly imprecise because each subgroup contains only a fraction of the overall information. A trial adequately powered for its primary treatment comparison may provide weak evidence about whether the treatment effect differs between age groups, sexes, institutions, or other subgroups.
This becomes particularly important when evaluating subgroup findings that were not prespecified.
Underpowering is not only a problem for nonsignificant results
Low-powered studies can also produce statistically significant findings. When an estimate happens to cross a significance threshold despite limited information, its magnitude may be unstable or imprecise. Publication and selective-reporting processes can further distort the collection of significant results that become visible in the literature.
Therefore, do not stop worrying about study information simply because P < 0.05. Inspect the confidence interval, event counts, design assumptions, and analytical process regardless of the significance label.
04 · A Practical Example
Recognizing When a Study Did Not Answer the Question It Was Designed For
Hypothetical Example
A trial recruits far fewer participants than planned
Suppose researchers planned a randomized trial to detect a 5-point improvement in an outcome. Their prespecified calculation required 240 analyzable participants for 90% power at their chosen significance level, and they planned to recruit 270 to allow for missing outcome data. Recruitment proves difficult, however, and only 142 participants contribute data to the primary analysis.
Read the original design. The study was explicitly planned around a 5-point target difference and 240 analyzable participants.
Compare planned and achieved information. Only 142 participants enter the primary analysis, substantially fewer than planned. The original 90% power statement therefore cannot simply be applied to the completed analysis.
Read the observed estimate. Suppose the estimated difference is 2 points with a 95% confidence interval from -4 to 8 points and P = 0.41.
Interpret the interval. The interval includes no effect, the prespecified 5-point target, and effects larger than that target. The study therefore does not distinguish adequately among several substantively different possibilities.
State the limitation precisely. Rather than saying "P = 0.41 proves there is no effect" or merely "the study was underpowered," explain that the study recruited substantially fewer analyzable participants than planned and produced an estimate too imprecise to exclude the target effect it was designed to detect.
This description tells the reader what the information problem actually means. "Underpowered" is useful shorthand only when you can connect it to the effect and inferential question that the study was supposed to address.
07 · A Quick Checklist
How to Check a Published Paper for Possible Underpowering
When assessing whether the study had enough information, check:
Which primary outcome and statistical analysis the sample-size calculation was based on.
What effect magnitude the researchers planned to detect and how that target was justified.
What significance level, target power, variability, event rate, and other assumptions were used where applicable.
Whether the planned number of analyzable participants or relevant events was actually achieved.
Whether attrition, missing data, rare outcomes, clustering, or exclusions substantially reduced available information.
How wide the confidence interval is and whether it still includes substantively important effects.
Whether secondary and subgroup analyses are being treated as though they had the same power as the primary analysis.
Whether claims of "no effect" are justified by precision rather than merely by P > 0.05.