Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Should You Look for When a Paper Reports Many Statistical Tests?

When a paper reports many statistical tests, individual P-values cannot always be interpreted in isolation. Examine why so many tests were performed, which analyses were primary, whether they were prespecified, and how multiplicity affected the conclusions.

390
When a Paper Reports Many Statistical Tests Guide 390 of 899
01 · The Question

What Changes When a Paper Tests Many Hypotheses?

You open a Results section and encounter P-values everywhere. The researchers compare several outcomes, multiple time points, numerous subgroups, alternative models, perhaps dozens of individual variables. Some results are statistically significant. Many are not.

The problem is not simply that the paper contains a lot of statistics. Multiple analyses can be entirely legitimate. A complex study may genuinely require several outcomes, comparisons, sensitivity analyses, or exploratory questions.

The critical issue is what those tests represent. Were they prespecified parts of a coherent analysis, or were many possibilities examined until something interesting appeared? Were all results reported, or mainly the favorable ones? And when several tests contribute to the same confirmatory claim, was multiplicity considered appropriately?

02 · The Short Answer

Do Not Count P-Values; Reconstruct the Analytical Strategy

In Brief

When a paper reports many statistical tests, identify the primary hypotheses and outcomes, distinguish prespecified analyses from exploratory or post hoc analyses, determine whether the tests belong to a common inferential family, and examine how the researchers addressed the increased opportunity for chance findings where multiplicity matters.

There is no rule that every collection of multiple tests requires the same correction. The appropriate response depends on the study's objectives and inferential structure. What should concern you most is a large, flexible analytical search presented as though every nominally significant result were an independent confirmatory finding.

03 · What You Need to Know

Many Tests Create More Opportunities to Find Something

Why multiplicity can become a statistical problem

Suppose 20 independent null hypotheses are all true and each is tested at a significance level of 0.05. The probability that any particular test produces a false positive is 5% under those assumptions. But the probability of obtaining at least one P < 0.05 somewhere among the 20 tests is much larger.

A Simple Illustration
1 - (1 - 0.05)20 ≈ 0.642
Under the simplifying assumptions that all 20 null hypotheses are true and the tests are independent, this is the probability of obtaining at least one nominally significant result when each test uses α = 0.05.
1 - 0.9520 ≈ 0.642, or about 64%. This does not mean that every real set of 20 tests has a 64% false-positive probability because tests are often correlated and the relevant error criterion depends on the inferential structure. It illustrates why the number and relationship of tests matter.

The concern is usually described as multiplicity or the multiple-testing problem. ICH statistical guidance identifies several common sources, including multiple primary variables, multiple treatment comparisons, repeated evaluation over time, and interim analyses. CONSORT 2025 likewise highlights multiple primary outcomes, time points, planned analyses, subgroups, and numerous secondary outcomes as settings in which multiplicity deserves attention.

First identify what the tests are trying to accomplish

Twenty tests do not automatically constitute one family of 20 hypotheses that requires a single universal correction. Statistical practice distinguishes among confirmatory hypotheses, secondary analyses, sensitivity analyses, exploratory investigations, safety analyses, and other purposes.

Ask what claim each test is supposed to support. If several tests collectively provide multiple opportunities to establish the same confirmatory conclusion, control of an appropriate error rate may be important. If analyses are explicitly exploratory, adjustment may or may not be the central issue, but the results should be interpreted as exploratory rather than presented as definitive discoveries.

That distinction is why mechanically counting every P-value in a paper can be misleading. You need the analytical architecture, not just the total.

Find the primary outcome and primary analysis

A well-designed confirmatory study should make its central question recognizable. In randomized trials, CONSORT 2025 calls for prespecified primary and secondary outcomes and notes that the primary outcome is normally the outcome of greatest importance used in the sample-size calculation.

When reading a paper with many results, locate the primary outcome and analysis before examining whichever P-value is smallest.

Question Why it matters
Which outcome was designated primary? Shows which result was intended to address the central confirmatory question
Which analyses were prespecified? Helps distinguish planned inference from analyses suggested by the observed data
How many primary hypotheses were there? Multiple opportunities to establish the main claim may create a multiplicity problem
Were multiple time points tested? Repeated opportunities to select a favorable time point can affect interpretation
Were many subgroups examined? Chance subgroup findings become increasingly plausible as analytical opportunities multiply
Were multiplicity procedures specified? Shows whether the inferential consequences were anticipated and addressed

Look for the protocol, registration, and statistical analysis plan

The final paper shows you what was reported. It may not show every analysis that was possible or attempted.

When available, compare the publication with a protocol, study registration, or statistical analysis plan. CONSORT 2025 recommends identifying which analyses were prespecified and which were post hoc, and explaining deviations from the planned analysis. This comparison can reveal whether outcomes, time points, models, or subgroup analyses changed after the data were available.

This matters because multiplicity is partly a problem of analytical opportunity. Ten transparently prespecified analyses are easier to interpret than an unknown number of analyses from which ten favorable results were selected.

Understand what multiple-testing procedures are trying to control

Different multiplicity procedures address different inferential goals. One common target is the family-wise error rate, the probability of making at least one false rejection within a defined family of hypotheses. Bonferroni adjustment is a simple method for controlling this error rate by using a more stringent threshold for each test.

Other procedures include Holm's sequential method and more specialized hierarchical or gatekeeping strategies. In settings involving large numbers of exploratory hypotheses, researchers may instead control quantities such as the false discovery rate.

The appropriate method depends on the scientific objective, dependence among tests, ordering of hypotheses, and what type of error the researchers are trying to control. "They should have used Bonferroni" is therefore not a universal appraisal rule.

No adjustment does not automatically make a paper invalid

Multiplicity does not imply that every P-value in every paper must be mathematically adjusted. Some analyses may address distinct questions, some may be descriptive, and some may be explicitly exploratory. In other settings, a prespecified hierarchy of hypotheses may determine which tests can be interpreted confirmatorily.

What matters is whether the statistical treatment and language of inference match the analytical purpose. CONSORT recommends reporting any methods used to account for multiplicity and also reporting when no such method was used, particularly when many analyses were performed.

For confirmatory clinical trials, ICH E9 advises that remaining multiplicity should be identified prospectively and that adjustment should be considered, with the chosen procedure or justification documented in the analysis plan.

Do not treat every significant secondary result as equivalent to the primary finding

Secondary outcomes can provide valuable evidence, but their interpretation depends on the study's inferential strategy. A paper may have one primary analysis and 25 secondary analyses. Finding P = 0.03 for one secondary endpoint after the primary analysis is null does not automatically rescue the study's main hypothesis.

Ask whether secondary analyses were prespecified, whether they were intended to support confirmatory claims, whether multiplicity was addressed where necessary, and whether the authors characterize them appropriately.

A useful warning sign is a Discussion section that quietly shifts emphasis away from a disappointing primary result toward whichever secondary result happened to cross 0.05.

Subgroups multiply analytical opportunities quickly

Suppose researchers examine treatment effects by sex, age, baseline severity, institution, prior treatment, socioeconomic status, and several biomarker categories. The number of potential comparisons can expand rapidly.

Subgroup appraisal also involves another common error: statistical significance in one subgroup and nonsignificance in another does not itself demonstrate that the subgroup effects differ. The appropriate question usually concerns an interaction or direct comparison between effects.

When a subgroup finding was not prespecified, multiplicity and data-driven analysis become especially relevant.

Many models can create hidden multiplicity

Multiplicity is not limited to a table containing 30 P-values. Researchers may try alternative covariate sets, outcome definitions, exclusion criteria, transformations, interaction terms, time windows, missing-data procedures, or model specifications.

If only the final preferred model appears in the paper, the reader may see one P-value while the researchers effectively had many opportunities to obtain it. This broader analytical flexibility is sometimes described as a researcher-degrees-of-freedom problem.

That is one reason recognizing selective reporting of significant analyses requires more than counting the tests visible in the final publication.

Effect sizes and confidence intervals remain essential

Multiplicity does not turn statistical interpretation into a contest among corrected P-values. Even when an appropriate multiple-testing procedure is used, you still need to know how large the effects are and how uncertain they remain.

A multiplicity-adjusted P-value below a threshold does not establish practical importance. Conversely, an estimate may be substantively interesting while its uncertainty remains too large for a firm conclusion.

Continue to interpret P-values alongside effect estimates and confidence intervals.

Large samples can make the situation look even more impressive

When a study combines many statistical tests with a very large dataset, numerous small associations may become statistically detectable. Some may be genuine but trivial, while others may reflect the opportunities created by multiplicity or analytical flexibility.

Therefore, when very large samples make tiny effects statistically significant, the number of analyses and the substantive magnitude of each effect both deserve scrutiny.

Watch Out

A paper may display only the successful end of a much larger analytical search. The number of tests you can count in the article is not necessarily the number of analyses that were attempted. Protocols, registrations, statistical analysis plans, supplements, and transparent reporting help reveal the difference.

04 · A Practical Example

What Happens When One Result Stands Out Among Many Tests?

Hypothetical Example

A study evaluates one intervention across many outcomes

Suppose researchers compare an educational intervention with usual teaching. They report one primary outcome, 12 secondary outcomes measured at two time points, and eight subgroup analyses. The prespecified primary outcome produces P = 0.24. Among the secondary analyses, one outcome at one time point produces P = 0.018, and the Discussion describes this as evidence that the intervention is effective.

Start with the primary question. The prespecified primary analysis did not provide conventional statistical evidence against its null hypothesis. That result should remain visible when interpreting the study.
Map the additional analyses. The paper contains many opportunities for favorable findings across secondary outcomes, time points, and subgroups.
Check prespecification. Determine whether the P = 0.018 analysis was specified in advance and what inferential status the statistical analysis plan gave it.
Check multiplicity. Find out whether the analysis plan used a correction, hierarchy, gatekeeping procedure, or another strategy, or explicitly treated secondary findings as exploratory.
Read the effect itself. Examine the magnitude and confidence interval rather than allowing P = 0.018 to carry the interpretation alone.
Match the conclusion to the evidence. If the finding emerged from a large collection of secondary analyses without an appropriate confirmatory strategy, it may reasonably generate a hypothesis, but presenting it as though it independently establishes the intervention's effectiveness would overstate the evidence.
05 · What Researchers Often Get Wrong

Common Mistakes When Appraising Multiple Statistical Tests

Misconception

Every Paper With Many Tests Must Use Bonferroni Correction

No. Multiplicity methods depend on the inferential structure and scientific objectives. Bonferroni is one option for particular families of hypotheses, not a universal requirement for every collection of analyses.

Misconception

If Five Results Have P < 0.05, All Five Are Confirmed Findings

Not necessarily. Their interpretation depends on how many hypotheses were examined, whether they were prespecified, whether they belong to a common inferential family, and how multiplicity and selective reporting were handled.

Misconception

Only the Tests Printed in the Paper Matter

The published analyses may be a subset of those attempted. Comparing the article with protocols, registrations, analysis plans, and supplements can reveal additional outcomes or analyses and help assess selective reporting.

Misconception

A Multiplicity-Adjusted Significant Result Must Be Important

Adjustment addresses a statistical error criterion, not substantive importance. Effect magnitude, precision, study validity, and real-world consequences still require separate evaluation.

Misconception

No Multiplicity Adjustment Means the Study Is Automatically Invalid

Not every analysis requires the same correction. The concern depends on what claims are being made, how hypotheses relate to one another, whether analyses are confirmatory or exploratory, and whether the inferential strategy was transparent and defensible.

06 · What This Means for You

Reconstruct the Path From Research Question to Reported Finding

When confronted with dozens of tests, do not attempt to judge each P-value in isolation. Work backward from the claims the authors make and identify which analyses are supposed to justify them.

A simple decision framework

If the paper has one clearly prespecified primary analysis and multiple supportive analyses
Give the primary analysis its intended inferential role and judge the others according to their prespecified purposes.
If several tests can independently establish the same confirmatory claim
Look for an appropriate multiplicity strategy and justification.
If many exploratory analyses are clearly identified as exploratory
Treat them as hypothesis-generating evidence and examine effect sizes, uncertainty, plausibility, and independent corroboration.
If the paper highlights isolated significant results from a large or poorly documented analytical search
Increase scrutiny for multiplicity, selective reporting, and data-driven interpretation.

The question is not simply "Did they correct for multiple comparisons?" A stronger appraisal asks what family of claims was being tested, how the analytical strategy was determined, and whether the strength of the conclusions matches that strategy.

07 · A Quick Checklist

When a Paper Contains Many Statistical Tests

Before accepting the significant findings, check:
Which outcome, hypothesis, and analysis were designated primary.
Which outcomes, time points, subgroup analyses, and models were prespecified and which were post hoc.
Whether several tests provide multiple opportunities to support the same confirmatory claim.
Whether the protocol, registration, or statistical analysis plan agrees with the analyses reported in the paper.
Whether an appropriate multiplicity procedure, hierarchy, or other inferential strategy was prespecified where necessary.
Whether secondary and exploratory findings are described with appropriately cautious language.
Whether effect estimates and confidence intervals support the substantive importance of highlighted findings.
Whether the Discussion disproportionately emphasizes isolated significant findings while minimizing null or contradictory analyses.
08 · Frequently Asked Questions

Questions About Multiple Statistical Testing

Why does performing many statistical tests increase false-positive concerns?

When multiple true null hypotheses are tested, each test creates an opportunity for a false rejection. Across a family of tests, the probability of at least one false positive can therefore exceed the nominal error rate of an individual test, depending on the number and dependence of the tests.

Does every study with multiple outcomes need a correction?

No universal correction applies to every study. The appropriate approach depends on which outcomes support confirmatory claims, how hypotheses are organized, the intended error criterion, and whether analyses are confirmatory or exploratory.

What is Bonferroni correction?

Bonferroni is a simple method for controlling the family-wise error rate. For a family of m hypotheses and desired family-wise level α, each hypothesis can be tested at α/m. It is broadly applicable but can be conservative, particularly when many hypotheses are tested.

Are exploratory analyses bad?

No. Exploration is a legitimate and valuable part of research. The problem arises when data-driven exploratory findings are presented as though they had the same evidential status as prespecified confirmatory tests.

Does preregistration solve multiple-testing problems?

Not by itself. Prespecification makes the intended analyses and hypotheses more transparent, but a prespecified design can still contain multiplicity that requires appropriate handling. It mainly helps distinguish planned analyses from those developed after seeing the data.

Should secondary outcomes be ignored if the primary outcome is nonsignificant?

No. Secondary outcomes can contain useful evidence. Their interpretation should reflect their prespecified role, the study's multiplicity strategy, effect estimates and uncertainty, and whether the claims are confirmatory or exploratory.

How can I tell whether more analyses were performed than the paper reports?

You may not be able to know completely. Compare the publication with the protocol, registration, statistical analysis plan, supplements, and other available reports. Discrepancies in outcomes, time points, subgroups, or model specifications can provide evidence of reporting changes.

09 · The Bottom Line

Many Tests Require Context, Not an Automatic Correction Formula

The Bottom Line

When a paper reports many statistical tests, determine which hypotheses were primary, which analyses were prespecified, how the tests relate to the claims being made, and whether multiplicity was handled appropriately for that inferential structure. Do not interpret isolated P-values as though the other analytical opportunities did not exist.

Multiple testing is not inherently improper, and no single correction applies universally. The central appraisal question is whether the analytical strategy was transparent and whether the authors' confidence in highlighted findings is proportionate to how those findings were generated.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes