Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What if Almost Every Study Is Underpowered?

When many studies are underpowered, a large literature may still provide surprisingly uncertain evidence. Learn how to recognize the pattern and interpret it without equating small samples with poor research.

758
What if Almost Every Study Is Underpowered? Guide 758 of 899
01 · The Question

What does it mean when an entire literature seems underpowered?

You read study after study and notice a recurring problem. Sample sizes are modest, estimates are imprecise, and many studies seem capable of detecting only fairly large effects. Some report statistically significant findings. Others report no statistically significant effect. Taken together, the literature looks substantial because there are many papers, yet each individual study contributes limited information.

What should you conclude? Is this merely a limitation to mention in your review, or does widespread low statistical power change what the literature can actually support?

The answer requires some care. A small sample is not automatically an underpowered sample, and statistical power cannot be judged from sample size alone. More importantly, once a study has been completed, calculating “observed power” from its observed effect is generally not a useful way to diagnose the problem. What matters is whether the study design could provide sufficiently informative evidence about effects that are scientifically or practically relevant.

02 · The Short Answer

Repeated low power can weaken an entire body of evidence

In Brief

If many studies have insufficient statistical power for effects that would matter, the literature may contain many results without providing correspondingly strong evidence.

Low power makes meaningful effects easier to miss and estimates less precise. When statistical significance influences which results become prominent or published, the significant effects that survive can also be unusually large. A review should therefore examine the information provided by the studies, not simply count significant and non-significant findings.

03 · What You Need to Know

How low statistical power changes the evidence you are reviewing

Underpowered does not simply mean “small sample”

Statistical power is the probability, under specified assumptions, that a statistical test will reject the null hypothesis when a particular non-null effect exists. Power depends on more than the number of observations. It also depends on the effect size being considered, variability or measurement precision, the statistical model and design, the significance threshold, and other features of the analysis.

That makes statements such as “studies with fewer than 100 participants are underpowered” generally indefensible without additional context. A study of 40 participants might provide considerable information about a very large, precisely measured effect. A study of 400 might still provide inadequate information about a small effect, a rare outcome, a noisy measure, an interaction, or a subgroup comparison.

Small sample Describes the number of observations relative to some context. By itself, it does not establish inadequate statistical power.
Low statistical power Means the design has a low probability of detecting a specified effect under specified assumptions.

Power is always power to detect something

Saying that a study has “30% power” is incomplete unless you know the effect size and assumptions used to obtain that figure. Power rises as the assumed effect becomes larger. Consequently, almost any finite study can be described as highly powered for a sufficiently large effect and poorly powered for a sufficiently small one.

This is why the effect size used in a power analysis deserves scrutiny. Ideally, it should reflect an effect that is plausible and meaningful for the research question. Lakens argues that sample-size justification should be tied to the inferential goal of the study, including consideration of the smallest effect size of interest, expected effects, desired precision, or the range of effects a design can detect. Merely reporting that a conventional power threshold was reached does not guarantee that the design was informative.

Low power makes true effects easier to miss

The most familiar consequence of low power is a greater probability of failing to obtain statistical significance when the effect specified in the power calculation is genuinely present. If many studies have low power, a literature can therefore accumulate numerous statistically non-significant results even when effects of interest exist.

This matters when reviewing a field because “not statistically significant” does not mean “no effect.” A non-significant estimate accompanied by a wide confidence interval may be compatible with no meaningful effect, a modest effect, or an important effect. In such cases, the study has not necessarily provided evidence of absence. It may simply provide insufficiently precise evidence to distinguish among those possibilities.

The significant findings can also be misleadingly large

Low power does not create only a problem of missed effects. When estimates are noisy and attention or publication depends on crossing a statistical-significance threshold, the estimates that happen to cross that threshold tend to be unusually large. Button and colleagues discussed this problem in relation to low-powered neuroscience research, while Gelman and Carlin describe the related risk of magnitude, or Type M, errors.

Imagine that an underlying effect is modest. Across repeated small studies, sampling variation will produce estimates above and below that effect. The unusually large estimates are more likely to reach a significance threshold. If statistically significant studies are preferentially published, cited, or emphasized, the visible literature can consequently exaggerate the apparent magnitude of the relationship.

Watch Out

Do not infer that every statistically significant result from a low-powered design is false. The defensible concern is that estimates selected because they reached statistical significance may be unstable or exaggerated, especially when combined with selective reporting or publication bias.

Many underpowered studies do not automatically become strong evidence by accumulation

It is tempting to reason that 20 small studies must collectively solve the problem of each study being weak. Sometimes synthesis genuinely can increase precision. Meta-analysis exists partly because combining compatible information across studies can estimate an effect more precisely than individual studies can.

But pooling is not statistical alchemy. Its interpretation depends on what evidence entered the synthesis. Selective publication, selective outcome reporting, heterogeneity, correlated samples, poor measurement, or other shared design limitations do not disappear because a meta-analysis produces a narrow confidence interval. If smaller studies systematically report larger effects, Cochrane guidance recommends examining possible small-study effects rather than assuming that pooling has neutralized them.

This is one reason a literature can become methodologically repetitive rather than genuinely cumulative. Twenty studies repeating the same informational limitation are not necessarily equivalent to a program of research in which later studies progressively reduce uncertainty.

Precision may be more informative than retrospective “observed power”

When reviewing completed studies, avoid mechanically calculating post hoc power by inserting each study's observed effect size into a power calculator. Hoenig and Heisey showed why this form of observed-power calculation is misleading: once the observed result is known, the calculation is largely a transformation of information already contained in the significance test.

Instead, examine what the completed study actually tells you. Effect estimates and confidence intervals reveal the magnitude and precision of the evidence. If a confidence interval remains compatible with substantially different conclusions, imprecision is itself an important finding. Where appropriate, sensitivity or design analyses based on externally justified plausible effects can also show what effects the original design was capable of detecting with reasonable probability.

Underpowering can interact with other recurring weaknesses

Power should not be evaluated in isolation. A literature may combine modest samples with noisy or poorly validated measures, which can further reduce informational value. Heavy reliance on convenience samples raises a different problem concerning generalizability. A field dominated by cross-sectional designs may estimate associations precisely while still being unable to resolve temporal or causal questions.

These limitations are conceptually distinct. A large sample does not repair weak measurement or an unsuitable design, just as an excellent design does not guarantee adequate precision. A good literature review keeps those problems separate before considering how they interact.

04 · A Practical Example

When 18 studies still leave the effect uncertain

Hypothetical Example

A literature full of modest studies

Suppose you are reviewing 18 studies examining whether a teaching intervention improves academic performance. Most studies compare relatively small groups. Few report an a priori sample-size justification, and the typical confidence intervals around estimated effects are wide.

What you observe Six studies report statistically significant improvements, while twelve do not.
What you should not conclude You should not simply declare that six studies “support” the intervention and twelve show that it “does not work.”
What you inspect instead Examine effect estimates, confidence intervals, sample-size justifications, design assumptions, and the size of effects the studies could reasonably detect.
What emerges Many non-significant studies remain compatible with effects that would matter educationally, while the significant studies tend to report some of the largest estimated effects.
How the synthesis changes The central conclusion becomes one of imprecision and uncertain effect magnitude rather than a vote count of six positive versus twelve negative studies.

The important insight is not that all 18 studies are useless. Each contributes information. The problem is that the literature may have repeatedly collected less information per study than was needed to discriminate among scientifically meaningful possibilities. If the same pattern continues, another similarly sized study may add one more paper without resolving much uncertainty.

05 · What Researchers Often Get Wrong

Common mistakes when interpreting underpowered literatures

Misconception

“The sample is small, so the study must be underpowered”

Sample size alone is insufficient. Power depends on the effect under consideration, variability, design, analysis, significance threshold, and other assumptions. Explain why the design appears insufficient for relevant effects rather than applying an arbitrary sample-size cutoff.

Misconception

“A non-significant finding means the effect is absent”

A non-significant result may be highly informative when its confidence interval rules out effects that matter. But when estimates are imprecise, the same result may remain compatible with meaningful effects. Interpret the estimate and its uncertainty rather than the significance label alone.

Misconception

“The significant studies prove low power was not a problem”

Obtaining significance does not retroactively make a design well powered. In noisy, low-powered settings, statistically significant estimates can be especially susceptible to magnitude exaggeration when significance acts as a selection filter.

Misconception

“I can calculate observed power from every published effect”

Post hoc power calculated from the observed effect is generally not an informative diagnostic for a completed study. Examine confidence intervals, effect-size precision, the original sample-size rationale, and, when useful, sensitivity analyses based on independently justified effects.

Misconception

“Enough small studies automatically solve the problem”

Compatible studies can collectively provide substantial information, particularly through appropriate meta-analysis. But accumulation does not automatically remove publication bias, shared design limitations, heterogeneity, or small-study effects. The relevant question is how much credible information the body of evidence contains.

Misconception

“Every study should have exactly 80% power”

Eighty percent is a widely used convention, not a universal law of research design. Appropriate error rates and sample sizes depend on the inferential goal, plausible effect sizes, costs of errors, design, and available resources. The substantive justification matters more than ritual compliance with a single threshold.

06 · What This Means for You

How should widespread low power change your literature review?

First, stop treating statistical significance as the unit of evidence. Extract effect estimates and uncertainty wherever possible. A literature containing mostly non-significant results can mean something very different when those estimates are precise than when their confidence intervals are broad.

Second, identify what effect sizes would actually matter for the research question. If most studies could reliably detect only effects considerably larger than plausible or substantively important ones, that is a meaningful limitation of the evidence base. Be explicit about the assumptions behind that judgment.

Third, look for the pattern across studies. Are estimates from smaller studies systematically larger? Are statistically significant results disproportionately prominent? Do larger or more precise studies produce smaller estimates? Such observations do not by themselves prove publication bias, but they may justify investigating small-study effects and missing evidence.

Finally, let the diagnosis affect your recommendation. If the literature has repeatedly asked essentially the same question with designs that provide limited precision, calling generically for “more studies” may reproduce the problem. The more useful recommendation may be for larger or otherwise more informative designs, better measurement, collaborative data collection, preregistered confirmatory studies, or designs explicitly justified against meaningful effect sizes. Sometimes the literature needs better research rather than simply more research.

A simple decision framework

If studies are small but estimates are sufficiently precise for the question
Do not label the literature underpowered merely because the samples look small.
If confidence intervals repeatedly include effects ranging from negligible to substantively important
Describe the evidence as imprecise and explain what conclusions remain unresolved.
If designs could detect only implausibly large effects with reasonable probability
Discuss limited sensitivity to realistic effects and justify the effect-size assumptions used for that assessment.
If small studies also show systematically larger effects
Investigate small-study effects and possible selective reporting rather than attributing the pattern to power alone.
If the same limitation recurs across most of the literature
Treat it as a property of the evidence base, not merely a footnote attached independently to each paper.
07 · A Quick Checklist

Before concluding that a literature is underpowered

Before writing the conclusion, check:
Whether the studies report prospective sample-size or power justifications and what assumptions those calculations used.
What effect sizes are plausible and substantively important for the research question.
Whether individual confidence intervals are sufficiently narrow to distinguish among conclusions that matter.
Whether you are equating small sample size with low power without considering design, variability, measurement, and analysis.
Whether non-significant findings are being incorrectly interpreted as evidence that no meaningful effect exists.
Whether statistically significant estimates from smaller studies appear systematically larger than estimates from larger or more precise studies.
Whether selective publication or outcome reporting could make the visible literature unrepresentative of all studies conducted.
Whether the cumulative evidence remains imprecise even after appropriate synthesis.
Whether your recommendation specifies how future studies should become more informative instead of merely requesting a larger number of studies.
08 · Frequently Asked Questions

Questions about underpowered studies in a literature review

What exactly is an underpowered study?

It is a study whose design has a relatively low probability of detecting a specified effect under specified assumptions. Because power depends on the effect size and design assumptions, a study is not meaningfully described as “underpowered” without clarifying what effect it was expected to detect.

Does a small sample always mean low statistical power?

No. Sample size is an important determinant of power, but power also depends on the effect size, variability, measurement, statistical design, analysis, significance threshold, and related assumptions. Small sample size alone does not establish inadequate power.

Can I calculate post hoc power for published studies?

You can perform calculations using a published design, but “observed power” based on the study's observed effect size is generally not a useful way to interpret the completed result. Effect estimates, confidence intervals, prospective sample-size justifications, and sensitivity analyses based on independently justified effect sizes are usually more informative.

Does low power increase false positives?

Low power is principally associated with a higher probability of missing specified non-null effects. The interpretation of significant findings becomes more complicated when low power is combined with significance-based selection, multiple testing, flexible analyses, or publication bias. Under those conditions, published significant estimates may be exaggerated or otherwise unrepresentative. Low power should not be described as mechanically increasing the nominal Type I error rate of a correctly specified test.

Can meta-analysis fix an underpowered literature?

It can improve precision when sufficiently comparable studies provide usable information. It cannot automatically correct publication bias, selective reporting, common methodological weaknesses, or all forms of heterogeneity. The quality and completeness of the evidence entering the synthesis still matter.

Should I exclude underpowered studies from my review?

Not automatically. Inclusion criteria should follow the review question and protocol rather than a blanket sample-size rule. Small or imprecise studies may still contribute information. Their limitations should instead influence risk-of-bias assessment where applicable, synthesis choices, sensitivity analyses, and the certainty attached to conclusions.

Is 80% power always the minimum acceptable level?

No. Eighty percent is a common convention rather than a universal methodological requirement. The appropriate design should be justified in relation to inferential goals, relevant effect sizes, acceptable error rates, precision, and practical constraints.

What should I recommend if most studies are underpowered?

Be specific about what would make future evidence more informative. Depending on the problem, that might mean larger samples, multicenter collaboration, more precise measurement, more efficient designs, justified effect-size targets, prospective sample-size planning, or confirmatory studies designed to distinguish among substantively meaningful effects.

09 · The Bottom Line

A large literature can still contain too little information

The Bottom Line

If most studies provide insufficient power or precision for effects that matter, the number of published papers can substantially overstate how much the field actually knows.

Do not diagnose the problem from sample size alone or from observed post hoc power. Examine plausible effect sizes, design assumptions, effect estimates, confidence intervals, and patterns across studies. When the same limitation recurs, the appropriate conclusion may be that the field needs studies designed to reduce uncertainty, not merely another study of the same size.

10 · Sources and Further Reading

Sources and further reading

Do you like this personality?
11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes