01 · The Question
What does it mean when an entire literature seems underpowered?
You read study after study and notice a recurring problem. Sample sizes are modest, estimates are imprecise, and many studies seem capable of detecting only fairly large effects. Some report statistically significant findings. Others report no statistically significant effect. Taken together, the literature looks substantial because there are many papers, yet each individual study contributes limited information.
What should you conclude? Is this merely a limitation to mention in your review, or does widespread low statistical power change what the literature can actually support?
The answer requires some care. A small sample is not automatically an underpowered sample, and statistical power cannot be judged from sample size alone. More importantly, once a study has been completed, calculating “observed power” from its observed effect is generally not a useful way to diagnose the problem. What matters is whether the study design could provide sufficiently informative evidence about effects that are scientifically or practically relevant.
02 · The Short Answer
Repeated low power can weaken an entire body of evidence
In Brief
If many studies have insufficient statistical power for effects that would matter, the literature may contain many results without providing correspondingly strong evidence.
Low power makes meaningful effects easier to miss and estimates less precise. When statistical significance influences which results become prominent or published, the significant effects that survive can also be unusually large. A review should therefore examine the information provided by the studies, not simply count significant and non-significant findings.
03 · What You Need to Know
How low statistical power changes the evidence you are reviewing
Underpowered does not simply mean “small sample”
Statistical power is the probability, under specified assumptions, that a statistical test will reject the null hypothesis when a particular non-null effect exists. Power depends on more than the number of observations. It also depends on the effect size being considered, variability or measurement precision, the statistical model and design, the significance threshold, and other features of the analysis.
That makes statements such as “studies with fewer than 100 participants are underpowered” generally indefensible without additional context. A study of 40 participants might provide considerable information about a very large, precisely measured effect. A study of 400 might still provide inadequate information about a small effect, a rare outcome, a noisy measure, an interaction, or a subgroup comparison.
Small sample
Describes the number of observations relative to some context. By itself, it does not establish inadequate statistical power.
Low statistical power
Means the design has a low probability of detecting a specified effect under specified assumptions.
Power is always power to detect something
Saying that a study has “30% power” is incomplete unless you know the effect size and assumptions used to obtain that figure. Power rises as the assumed effect becomes larger. Consequently, almost any finite study can be described as highly powered for a sufficiently large effect and poorly powered for a sufficiently small one.
This is why the effect size used in a power analysis deserves scrutiny. Ideally, it should reflect an effect that is plausible and meaningful for the research question. Lakens argues that sample-size justification should be tied to the inferential goal of the study, including consideration of the smallest effect size of interest, expected effects, desired precision, or the range of effects a design can detect. Merely reporting that a conventional power threshold was reached does not guarantee that the design was informative.
Low power makes true effects easier to miss
The most familiar consequence of low power is a greater probability of failing to obtain statistical significance when the effect specified in the power calculation is genuinely present. If many studies have low power, a literature can therefore accumulate numerous statistically non-significant results even when effects of interest exist.
This matters when reviewing a field because “not statistically significant” does not mean “no effect.” A non-significant estimate accompanied by a wide confidence interval may be compatible with no meaningful effect, a modest effect, or an important effect. In such cases, the study has not necessarily provided evidence of absence. It may simply provide insufficiently precise evidence to distinguish among those possibilities.
The significant findings can also be misleadingly large
Low power does not create only a problem of missed effects. When estimates are noisy and attention or publication depends on crossing a statistical-significance threshold, the estimates that happen to cross that threshold tend to be unusually large. Button and colleagues discussed this problem in relation to low-powered neuroscience research, while Gelman and Carlin describe the related risk of magnitude, or Type M, errors.
Imagine that an underlying effect is modest. Across repeated small studies, sampling variation will produce estimates above and below that effect. The unusually large estimates are more likely to reach a significance threshold. If statistically significant studies are preferentially published, cited, or emphasized, the visible literature can consequently exaggerate the apparent magnitude of the relationship.
Watch Out
Do not infer that every statistically significant result from a low-powered design is false. The defensible concern is that estimates selected because they reached statistical significance may be unstable or exaggerated, especially when combined with selective reporting or publication bias.
Many underpowered studies do not automatically become strong evidence by accumulation
It is tempting to reason that 20 small studies must collectively solve the problem of each study being weak. Sometimes synthesis genuinely can increase precision. Meta-analysis exists partly because combining compatible information across studies can estimate an effect more precisely than individual studies can.
But pooling is not statistical alchemy. Its interpretation depends on what evidence entered the synthesis. Selective publication, selective outcome reporting, heterogeneity, correlated samples, poor measurement, or other shared design limitations do not disappear because a meta-analysis produces a narrow confidence interval. If smaller studies systematically report larger effects, Cochrane guidance recommends examining possible small-study effects rather than assuming that pooling has neutralized them.
This is one reason a literature can become methodologically repetitive rather than genuinely cumulative . Twenty studies repeating the same informational limitation are not necessarily equivalent to a program of research in which later studies progressively reduce uncertainty.
Precision may be more informative than retrospective “observed power”
When reviewing completed studies, avoid mechanically calculating post hoc power by inserting each study's observed effect size into a power calculator. Hoenig and Heisey showed why this form of observed-power calculation is misleading: once the observed result is known, the calculation is largely a transformation of information already contained in the significance test.
Instead, examine what the completed study actually tells you. Effect estimates and confidence intervals reveal the magnitude and precision of the evidence. If a confidence interval remains compatible with substantially different conclusions, imprecision is itself an important finding. Where appropriate, sensitivity or design analyses based on externally justified plausible effects can also show what effects the original design was capable of detecting with reasonable probability.
Underpowering can interact with other recurring weaknesses
Power should not be evaluated in isolation. A literature may combine modest samples with noisy or poorly validated measures , which can further reduce informational value. Heavy reliance on convenience samples raises a different problem concerning generalizability. A field dominated by cross-sectional designs may estimate associations precisely while still being unable to resolve temporal or causal questions.
These limitations are conceptually distinct. A large sample does not repair weak measurement or an unsuitable design, just as an excellent design does not guarantee adequate precision. A good literature review keeps those problems separate before considering how they interact.
04 · A Practical Example
When 18 studies still leave the effect uncertain
Hypothetical Example
A literature full of modest studies
Suppose you are reviewing 18 studies examining whether a teaching intervention improves academic performance. Most studies compare relatively small groups. Few report an a priori sample-size justification, and the typical confidence intervals around estimated effects are wide.
What you observe
Six studies report statistically significant improvements, while twelve do not.
What you should not conclude
You should not simply declare that six studies “support” the intervention and twelve show that it “does not work.”
What you inspect instead
Examine effect estimates, confidence intervals, sample-size justifications, design assumptions, and the size of effects the studies could reasonably detect.
What emerges
Many non-significant studies remain compatible with effects that would matter educationally, while the significant studies tend to report some of the largest estimated effects.
How the synthesis changes
The central conclusion becomes one of imprecision and uncertain effect magnitude rather than a vote count of six positive versus twelve negative studies.
The important insight is not that all 18 studies are useless. Each contributes information. The problem is that the literature may have repeatedly collected less information per study than was needed to discriminate among scientifically meaningful possibilities. If the same pattern continues, another similarly sized study may add one more paper without resolving much uncertainty.
06 · What This Means for You
How should widespread low power change your literature review?
First, stop treating statistical significance as the unit of evidence. Extract effect estimates and uncertainty wherever possible. A literature containing mostly non-significant results can mean something very different when those estimates are precise than when their confidence intervals are broad.
Second, identify what effect sizes would actually matter for the research question. If most studies could reliably detect only effects considerably larger than plausible or substantively important ones, that is a meaningful limitation of the evidence base. Be explicit about the assumptions behind that judgment.
Third, look for the pattern across studies. Are estimates from smaller studies systematically larger? Are statistically significant results disproportionately prominent? Do larger or more precise studies produce smaller estimates? Such observations do not by themselves prove publication bias, but they may justify investigating small-study effects and missing evidence.
Finally, let the diagnosis affect your recommendation. If the literature has repeatedly asked essentially the same question with designs that provide limited precision, calling generically for “more studies” may reproduce the problem. The more useful recommendation may be for larger or otherwise more informative designs, better measurement, collaborative data collection, preregistered confirmatory studies, or designs explicitly justified against meaningful effect sizes. Sometimes the literature needs better research rather than simply more research .
A simple decision framework
If studies are small but estimates are sufficiently precise for the question
Do not label the literature underpowered merely because the samples look small.
If confidence intervals repeatedly include effects ranging from negligible to substantively important
Describe the evidence as imprecise and explain what conclusions remain unresolved.
If designs could detect only implausibly large effects with reasonable probability
Discuss limited sensitivity to realistic effects and justify the effect-size assumptions used for that assessment.
If small studies also show systematically larger effects
Investigate small-study effects and possible selective reporting rather than attributing the pattern to power alone.
If the same limitation recurs across most of the literature
Treat it as a property of the evidence base, not merely a footnote attached independently to each paper.
07 · A Quick Checklist
Before concluding that a literature is underpowered
Before writing the conclusion, check:
Whether the studies report prospective sample-size or power justifications and what assumptions those calculations used.
What effect sizes are plausible and substantively important for the research question.
Whether individual confidence intervals are sufficiently narrow to distinguish among conclusions that matter.
Whether you are equating small sample size with low power without considering design, variability, measurement, and analysis.
Whether non-significant findings are being incorrectly interpreted as evidence that no meaningful effect exists.
Whether statistically significant estimates from smaller studies appear systematically larger than estimates from larger or more precise studies.
Whether selective publication or outcome reporting could make the visible literature unrepresentative of all studies conducted.
Whether the cumulative evidence remains imprecise even after appropriate synthesis.
Whether your recommendation specifies how future studies should become more informative instead of merely requesting a larger number of studies.
09 · The Bottom Line
A large literature can still contain too little information
The Bottom Line
If most studies provide insufficient power or precision for effects that matter, the number of published papers can substantially overstate how much the field actually knows.
Do not diagnose the problem from sample size alone or from observed post hoc power. Examine plausible effect sizes, design assumptions, effect estimates, confidence intervals, and patterns across studies. When the same limitation recurs, the appropriate conclusion may be that the field needs studies designed to reduce uncertainty, not merely another study of the same size.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation