01 · The Question
What Changes When a Paper Tests Many Hypotheses?
You open a Results section and encounter P-values everywhere. The researchers compare several outcomes, multiple time points, numerous subgroups, alternative models, perhaps dozens of individual variables. Some results are statistically significant. Many are not.
The problem is not simply that the paper contains a lot of statistics. Multiple analyses can be entirely legitimate. A complex study may genuinely require several outcomes, comparisons, sensitivity analyses, or exploratory questions.
The critical issue is what those tests represent. Were they prespecified parts of a coherent analysis, or were many possibilities examined until something interesting appeared? Were all results reported, or mainly the favorable ones? And when several tests contribute to the same confirmatory claim, was multiplicity considered appropriately?
03 · What You Need to Know
Many Tests Create More Opportunities to Find Something
Why multiplicity can become a statistical problem
Suppose 20 independent null hypotheses are all true and each is tested at a significance level of 0.05. The probability that any particular test produces a false positive is 5% under those assumptions. But the probability of obtaining at least one P < 0.05 somewhere among the 20 tests is much larger.
The concern is usually described as multiplicity or the multiple-testing problem. ICH statistical guidance identifies several common sources, including multiple primary variables, multiple treatment comparisons, repeated evaluation over time, and interim analyses. CONSORT 2025 likewise highlights multiple primary outcomes, time points, planned analyses, subgroups, and numerous secondary outcomes as settings in which multiplicity deserves attention.
First identify what the tests are trying to accomplish
Twenty tests do not automatically constitute one family of 20 hypotheses that requires a single universal correction. Statistical practice distinguishes among confirmatory hypotheses, secondary analyses, sensitivity analyses, exploratory investigations, safety analyses, and other purposes.
Ask what claim each test is supposed to support. If several tests collectively provide multiple opportunities to establish the same confirmatory conclusion, control of an appropriate error rate may be important. If analyses are explicitly exploratory, adjustment may or may not be the central issue, but the results should be interpreted as exploratory rather than presented as definitive discoveries.
That distinction is why mechanically counting every P-value in a paper can be misleading. You need the analytical architecture, not just the total.
Find the primary outcome and primary analysis
A well-designed confirmatory study should make its central question recognizable. In randomized trials, CONSORT 2025 calls for prespecified primary and secondary outcomes and notes that the primary outcome is normally the outcome of greatest importance used in the sample-size calculation.
When reading a paper with many results, locate the primary outcome and analysis before examining whichever P-value is smallest.
| Question |
Why it matters |
| Which outcome was designated primary? |
Shows which result was intended to address the central confirmatory question |
| Which analyses were prespecified? |
Helps distinguish planned inference from analyses suggested by the observed data |
| How many primary hypotheses were there? |
Multiple opportunities to establish the main claim may create a multiplicity problem |
| Were multiple time points tested? |
Repeated opportunities to select a favorable time point can affect interpretation |
| Were many subgroups examined? |
Chance subgroup findings become increasingly plausible as analytical opportunities multiply |
| Were multiplicity procedures specified? |
Shows whether the inferential consequences were anticipated and addressed |
Look for the protocol, registration, and statistical analysis plan
The final paper shows you what was reported. It may not show every analysis that was possible or attempted.
When available, compare the publication with a protocol, study registration, or statistical analysis plan. CONSORT 2025 recommends identifying which analyses were prespecified and which were post hoc, and explaining deviations from the planned analysis. This comparison can reveal whether outcomes, time points, models, or subgroup analyses changed after the data were available.
This matters because multiplicity is partly a problem of analytical opportunity. Ten transparently prespecified analyses are easier to interpret than an unknown number of analyses from which ten favorable results were selected.
Understand what multiple-testing procedures are trying to control
Different multiplicity procedures address different inferential goals. One common target is the family-wise error rate, the probability of making at least one false rejection within a defined family of hypotheses. Bonferroni adjustment is a simple method for controlling this error rate by using a more stringent threshold for each test.
Other procedures include Holm's sequential method and more specialized hierarchical or gatekeeping strategies. In settings involving large numbers of exploratory hypotheses, researchers may instead control quantities such as the false discovery rate.
The appropriate method depends on the scientific objective, dependence among tests, ordering of hypotheses, and what type of error the researchers are trying to control. "They should have used Bonferroni" is therefore not a universal appraisal rule.
No adjustment does not automatically make a paper invalid
Multiplicity does not imply that every P-value in every paper must be mathematically adjusted. Some analyses may address distinct questions, some may be descriptive, and some may be explicitly exploratory. In other settings, a prespecified hierarchy of hypotheses may determine which tests can be interpreted confirmatorily.
What matters is whether the statistical treatment and language of inference match the analytical purpose. CONSORT recommends reporting any methods used to account for multiplicity and also reporting when no such method was used, particularly when many analyses were performed.
For confirmatory clinical trials, ICH E9 advises that remaining multiplicity should be identified prospectively and that adjustment should be considered, with the chosen procedure or justification documented in the analysis plan.
Do not treat every significant secondary result as equivalent to the primary finding
Secondary outcomes can provide valuable evidence, but their interpretation depends on the study's inferential strategy. A paper may have one primary analysis and 25 secondary analyses. Finding P = 0.03 for one secondary endpoint after the primary analysis is null does not automatically rescue the study's main hypothesis.
Ask whether secondary analyses were prespecified, whether they were intended to support confirmatory claims, whether multiplicity was addressed where necessary, and whether the authors characterize them appropriately.
A useful warning sign is a Discussion section that quietly shifts emphasis away from a disappointing primary result toward whichever secondary result happened to cross 0.05.
Subgroups multiply analytical opportunities quickly
Suppose researchers examine treatment effects by sex, age, baseline severity, institution, prior treatment, socioeconomic status, and several biomarker categories. The number of potential comparisons can expand rapidly.
Subgroup appraisal also involves another common error: statistical significance in one subgroup and nonsignificance in another does not itself demonstrate that the subgroup effects differ. The appropriate question usually concerns an interaction or direct comparison between effects.
When a subgroup finding was not prespecified, multiplicity and data-driven analysis become especially relevant.
Many models can create hidden multiplicity
Multiplicity is not limited to a table containing 30 P-values. Researchers may try alternative covariate sets, outcome definitions, exclusion criteria, transformations, interaction terms, time windows, missing-data procedures, or model specifications.
If only the final preferred model appears in the paper, the reader may see one P-value while the researchers effectively had many opportunities to obtain it. This broader analytical flexibility is sometimes described as a researcher-degrees-of-freedom problem.
That is one reason recognizing selective reporting of significant analyses requires more than counting the tests visible in the final publication.
Effect sizes and confidence intervals remain essential
Multiplicity does not turn statistical interpretation into a contest among corrected P-values. Even when an appropriate multiple-testing procedure is used, you still need to know how large the effects are and how uncertain they remain.
A multiplicity-adjusted P-value below a threshold does not establish practical importance. Conversely, an estimate may be substantively interesting while its uncertainty remains too large for a firm conclusion.
Continue to interpret P-values alongside effect estimates and confidence intervals.
Large samples can make the situation look even more impressive
When a study combines many statistical tests with a very large dataset, numerous small associations may become statistically detectable. Some may be genuine but trivial, while others may reflect the opportunities created by multiplicity or analytical flexibility.
Therefore, when very large samples make tiny effects statistically significant, the number of analyses and the substantive magnitude of each effect both deserve scrutiny.
Watch Out
A paper may display only the successful end of a much larger analytical search. The number of tests you can count in the article is not necessarily the number of analyses that were attempted. Protocols, registrations, statistical analysis plans, supplements, and transparent reporting help reveal the difference.
04 · A Practical Example
What Happens When One Result Stands Out Among Many Tests?
Hypothetical Example
A study evaluates one intervention across many outcomes
Suppose researchers compare an educational intervention with usual teaching. They report one primary outcome, 12 secondary outcomes measured at two time points, and eight subgroup analyses. The prespecified primary outcome produces P = 0.24. Among the secondary analyses, one outcome at one time point produces P = 0.018, and the Discussion describes this as evidence that the intervention is effective.
Start with the primary question. The prespecified primary analysis did not provide conventional statistical evidence against its null hypothesis. That result should remain visible when interpreting the study.
Map the additional analyses. The paper contains many opportunities for favorable findings across secondary outcomes, time points, and subgroups.
Check prespecification. Determine whether the P = 0.018 analysis was specified in advance and what inferential status the statistical analysis plan gave it.
Check multiplicity. Find out whether the analysis plan used a correction, hierarchy, gatekeeping procedure, or another strategy, or explicitly treated secondary findings as exploratory.
Read the effect itself. Examine the magnitude and confidence interval rather than allowing P = 0.018 to carry the interpretation alone.
Match the conclusion to the evidence. If the finding emerged from a large collection of secondary analyses without an appropriate confirmatory strategy, it may reasonably generate a hypothesis, but presenting it as though it independently establishes the intervention's effectiveness would overstate the evidence.