01 · The Question
What If the Reported Analysis Was Only One of Many That Were Tried?
A paper reports that an intervention produced a statistically significant effect after adjusting for several covariates. The analysis looks reasonable. But what if the researchers also ran the analysis without adjustment, tried several combinations of covariates, used alternative definitions of the outcome, and examined multiple subgroups?
If only the analysis producing the most favorable result appears in the paper, the reported result no longer represents a neutral analytical choice. It has survived a selection process that the reader cannot see.
This is the central problem of selective analysis reporting.
03 · What You Need to Know
Why Analytical Flexibility Can Become Reporting Bias
Most datasets permit more than one analytical choice
Real research rarely presents researchers with a single inevitable statistical analysis. Investigators may need to decide how to handle missing data, which covariates to include, whether to analyze change scores or post-intervention scores, which transformations to apply, how to define an analytical population, whether to combine categories, or which subgroup comparisons to examine.
Many of these choices can be scientifically defensible. Analytical flexibility is therefore not inherently a methodological flaw.
The problem arises when the findings influence which analytical version becomes the reported result.
The key distinction is analysis choice before versus after seeing the results
Suppose two analytical approaches are both plausible. If researchers specify one in advance because it best matches the design and estimand of interest, the existence of an alternative does not by itself create selective reporting bias.
If researchers examine both results and then report only the analysis producing the stronger or more favorable finding, selection has become result-dependent. Cochrane describes this as bias in selection of the reported result when a result is selected, based on the findings, from multiple eligible analyses.
Analytical flexibility
More than one analytical approach is possible. This is common and not inherently biased.
Selective analysis reporting
The analysis that becomes visible is selected from alternatives because of the direction, magnitude, statistical significance, or other characteristics of the resulting estimate.
A reasonable-looking analysis can still be selectively reported
This is one reason selective analysis reporting is difficult to detect from a paper alone. The reported model does not need to be obviously incorrect.
Imagine that five defensible models were fitted and produced p-values of 0.18, 0.11, 0.07, 0.04, and 0.03. A paper reporting only the fifth model may contain a mathematically valid analysis. Yet the evidential meaning of p = 0.03 is different if that model was prespecified than if it was selected because it happened to cross a conventional threshold after several alternatives were examined.
The hidden multiplicity matters because readers usually interpret a reported analysis as though the analytical pathway leading to it was determined by the research question and design rather than by repeated inspection of the results.
Selective analysis reporting can take many forms
Cochrane identifies several circumstances in which multiple eligible analyses can generate different estimates. Examples include choosing between change scores and post-intervention scores, reporting adjusted rather than unadjusted estimates, adjusting for different sets of prognostic factors, converting continuous outcomes into categories using alternative cut-points, and selecting among alternative composite outcomes.
Analytical decision
Possible alternatives
How selective reporting can arise
Covariate adjustment
Unadjusted model or several adjusted models
Only the model producing the preferred estimate is emphasized
Missing data
Complete-case analysis or alternative assumptions and methods
The treatment producing the more favorable result is selected
Outcome coding
Continuous score or several categorical cut-points
The coding that yields the strongest association is reported
Analysis population
Intention-to-treat, per-protocol, or other eligible populations
The population yielding the preferred result receives emphasis
Subgroups
Several participant characteristics or thresholds
Only striking subgroup findings are highlighted
Empirical research has found discrepancies between planned and reported analyses
Dwan and colleagues systematically reviewed 22 cohort studies covering 3,140 randomized controlled trials and compared analytical information across sources such as protocols, registries, and publications. They found discrepancies involving statistical analyses, handling of missing data, adjusted versus unadjusted analyses, continuous data, composite outcomes, and subgroup analyses.
The discrepancy rates varied substantially among the included investigations. For example, studies examining adjusted versus unadjusted analyses reported discrepancy rates ranging from 46% to 82%, while studies examining subgroup analyses reported rates from 61% to 100%.
Those numbers should not be interpreted as the prevalence of biased analysis selection across all research. The review explicitly noted that discrepancies do not by themselves establish why an analysis changed, and the available evidence was insufficient to distinguish reliably among bias, error, and legitimate changes in many cases. The important finding is that analytical discrepancies were common enough to make transparency about planned and reported analyses consequential.
Selective analysis reporting is different from selective outcome reporting
The two problems are closely related but conceptually separable. With selective outcome reporting , the selection concerns which outcomes or outcome measurements become available. With selective analysis reporting, the outcome may remain the same while the statistical treatment of its data changes.
For example, imagine that anxiety is measured as planned. Omitting anxiety entirely because its result is uninteresting is an outcome-reporting problem. Reporting anxiety but choosing, from several models, the covariate adjustment that produces the most favorable result is an analysis-reporting problem.
Both mechanisms can operate in the same study.
Outcome switching is another neighboring problem
Selective analysis reporting should also be distinguished from outcome switching . Changing which outcome is designated as primary concerns the role assigned to an outcome. Selective analysis reporting concerns which analytical version of a result becomes visible.
A study can therefore retain exactly the same primary outcome while still being vulnerable to selective analysis reporting.
Why this can distort an entire literature
If result-dependent analytical selection occurs repeatedly across studies, the consequences accumulate. Published papers may contain a disproportionate share of estimates that happen to be larger, statistically significant, or supportive of investigators' expectations.
A later reviewer then encounters these selected results as though each were simply the study's result. The synthesis may therefore inherit selection processes that occurred before the papers ever reached the review.
This is conceptually related to publication bias , but the filtering occurs within the analytical results of studies rather than solely through whether entire studies become available.
Watch Out
A sophisticated statistical model is not evidence that the model was selected independently of its result. When several reasonable analyses were possible, the analytical history and prespecification may matter as much as the apparent sophistication of the final model.
Prespecification helps, but it does not forbid legitimate changes
A protocol or statistical analysis plan can establish which analyses were intended before investigators had access to the relevant results. This makes later deviations visible and reduces ambiguity about whether choices were data-driven.
Prespecification should not be treated as a prohibition against improving an analysis. Researchers sometimes discover errors in a planned method, encounter unexpected data problems, or identify a better analytical approach for defensible reasons. The more informative practice is to preserve the prespecified analysis, document deviations, explain why they occurred, and distinguish planned from exploratory work.
Transparency changes the interpretation. An openly labeled sensitivity or exploratory analysis is not equivalent to quietly replacing a prespecified analysis after seeing which version produces the preferred result.
06 · What This Means for You
Evaluate the Analytical Path, Not Just the Final Model
When reading a paper, you usually cannot reconstruct every analysis the researchers considered. You can, however, ask whether the reported analysis appears to have been specified before the results were known and whether important analytical alternatives are acknowledged.
A simple decision framework
If a statistical analysis plan or protocol is available
Compare the planned analytical approach with the one actually reported.
If the reported analysis differs from the original plan
Look for a dated explanation and determine whether the change could have occurred after relevant results were known.
If several analytical choices are equally plausible
Look for sensitivity analyses showing whether the substantive conclusion survives reasonable alternatives.
If an unexpected subgroup, cut-point, or adjusted model drives the conclusion
Treat its confirmatory status cautiously unless the choice was prespecified or independently replicated.
For your own work, make important analytical decisions before examining the relevant results whenever feasible. A dated statistical analysis plan can be valuable even outside fields where formal preregistration is customary. When plans must change, retain the original specification and explain the change rather than rewriting the analytical history.
Exploratory analysis remains useful. The goal is not to eliminate exploration but to identify it honestly. Readers can interpret a post hoc finding appropriately when they know that it emerged from exploration rather than from a single prespecified confirmatory test.
07 · A Quick Checklist
How to Check for Selective Analysis Reporting
When evaluating a reported analysis, check:
Whether a protocol, preregistration, or statistical analysis plan specifies the analysis in advance.
Whether the reported model matches the prespecified model and analytical population.
Whether covariates, transformations, cut-points, subgroup definitions, and missing-data methods were chosen before the relevant results were examined.
Whether deviations from the planned analysis are disclosed and justified.
Whether reasonable alternative analyses produce materially different conclusions.
Whether surprising subgroup or model-specific findings are presented as exploratory when appropriate.
Whether the paper reports enough analytical detail for readers to understand how the final result was obtained.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation