Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Can Selective Analysis Reporting Distort the Literature?

The same dataset can often be analyzed in several defensible ways. Selective analysis reporting becomes a problem when the version that gets reported is chosen because it produced the preferred result.

493
How Selective Analysis Reporting Distorts Evidence Guide 493 of 899
01 · The Question

What If the Reported Analysis Was Only One of Many That Were Tried?

A paper reports that an intervention produced a statistically significant effect after adjusting for several covariates. The analysis looks reasonable. But what if the researchers also ran the analysis without adjustment, tried several combinations of covariates, used alternative definitions of the outcome, and examined multiple subgroups?

If only the analysis producing the most favorable result appears in the paper, the reported result no longer represents a neutral analytical choice. It has survived a selection process that the reader cannot see.

This is the central problem of selective analysis reporting.

02 · The Short Answer

The Reported Analysis May Be a Selected Version of the Result

In Brief

Selective analysis reporting can distort research when investigators conduct multiple plausible analyses but report a subset because those analyses produced more favorable, statistically significant, larger, or otherwise preferred results.

The reported analysis may be technically legitimate when considered by itself. The bias arises because readers cannot see the alternative analyses that were available, or that were actually conducted, and therefore cannot account for the result-dependent selection process.

03 · What You Need to Know

Why Analytical Flexibility Can Become Reporting Bias

Most datasets permit more than one analytical choice

Real research rarely presents researchers with a single inevitable statistical analysis. Investigators may need to decide how to handle missing data, which covariates to include, whether to analyze change scores or post-intervention scores, which transformations to apply, how to define an analytical population, whether to combine categories, or which subgroup comparisons to examine.

Many of these choices can be scientifically defensible. Analytical flexibility is therefore not inherently a methodological flaw.

The problem arises when the findings influence which analytical version becomes the reported result.

The key distinction is analysis choice before versus after seeing the results

Suppose two analytical approaches are both plausible. If researchers specify one in advance because it best matches the design and estimand of interest, the existence of an alternative does not by itself create selective reporting bias.

If researchers examine both results and then report only the analysis producing the stronger or more favorable finding, selection has become result-dependent. Cochrane describes this as bias in selection of the reported result when a result is selected, based on the findings, from multiple eligible analyses.

Analytical flexibility More than one analytical approach is possible. This is common and not inherently biased.
Selective analysis reporting The analysis that becomes visible is selected from alternatives because of the direction, magnitude, statistical significance, or other characteristics of the resulting estimate.

A reasonable-looking analysis can still be selectively reported

This is one reason selective analysis reporting is difficult to detect from a paper alone. The reported model does not need to be obviously incorrect.

Imagine that five defensible models were fitted and produced p-values of 0.18, 0.11, 0.07, 0.04, and 0.03. A paper reporting only the fifth model may contain a mathematically valid analysis. Yet the evidential meaning of p = 0.03 is different if that model was prespecified than if it was selected because it happened to cross a conventional threshold after several alternatives were examined.

The hidden multiplicity matters because readers usually interpret a reported analysis as though the analytical pathway leading to it was determined by the research question and design rather than by repeated inspection of the results.

Selective analysis reporting can take many forms

Cochrane identifies several circumstances in which multiple eligible analyses can generate different estimates. Examples include choosing between change scores and post-intervention scores, reporting adjusted rather than unadjusted estimates, adjusting for different sets of prognostic factors, converting continuous outcomes into categories using alternative cut-points, and selecting among alternative composite outcomes.

Analytical decision Possible alternatives How selective reporting can arise
Covariate adjustment Unadjusted model or several adjusted models Only the model producing the preferred estimate is emphasized
Missing data Complete-case analysis or alternative assumptions and methods The treatment producing the more favorable result is selected
Outcome coding Continuous score or several categorical cut-points The coding that yields the strongest association is reported
Analysis population Intention-to-treat, per-protocol, or other eligible populations The population yielding the preferred result receives emphasis
Subgroups Several participant characteristics or thresholds Only striking subgroup findings are highlighted

Empirical research has found discrepancies between planned and reported analyses

Dwan and colleagues systematically reviewed 22 cohort studies covering 3,140 randomized controlled trials and compared analytical information across sources such as protocols, registries, and publications. They found discrepancies involving statistical analyses, handling of missing data, adjusted versus unadjusted analyses, continuous data, composite outcomes, and subgroup analyses.

The discrepancy rates varied substantially among the included investigations. For example, studies examining adjusted versus unadjusted analyses reported discrepancy rates ranging from 46% to 82%, while studies examining subgroup analyses reported rates from 61% to 100%.

Those numbers should not be interpreted as the prevalence of biased analysis selection across all research. The review explicitly noted that discrepancies do not by themselves establish why an analysis changed, and the available evidence was insufficient to distinguish reliably among bias, error, and legitimate changes in many cases. The important finding is that analytical discrepancies were common enough to make transparency about planned and reported analyses consequential.

Selective analysis reporting is different from selective outcome reporting

The two problems are closely related but conceptually separable. With selective outcome reporting, the selection concerns which outcomes or outcome measurements become available. With selective analysis reporting, the outcome may remain the same while the statistical treatment of its data changes.

For example, imagine that anxiety is measured as planned. Omitting anxiety entirely because its result is uninteresting is an outcome-reporting problem. Reporting anxiety but choosing, from several models, the covariate adjustment that produces the most favorable result is an analysis-reporting problem.

Both mechanisms can operate in the same study.

Outcome switching is another neighboring problem

Selective analysis reporting should also be distinguished from outcome switching. Changing which outcome is designated as primary concerns the role assigned to an outcome. Selective analysis reporting concerns which analytical version of a result becomes visible.

A study can therefore retain exactly the same primary outcome while still being vulnerable to selective analysis reporting.

Why this can distort an entire literature

If result-dependent analytical selection occurs repeatedly across studies, the consequences accumulate. Published papers may contain a disproportionate share of estimates that happen to be larger, statistically significant, or supportive of investigators' expectations.

A later reviewer then encounters these selected results as though each were simply the study's result. The synthesis may therefore inherit selection processes that occurred before the papers ever reached the review.

This is conceptually related to publication bias, but the filtering occurs within the analytical results of studies rather than solely through whether entire studies become available.

Watch Out

A sophisticated statistical model is not evidence that the model was selected independently of its result. When several reasonable analyses were possible, the analytical history and prespecification may matter as much as the apparent sophistication of the final model.

Prespecification helps, but it does not forbid legitimate changes

A protocol or statistical analysis plan can establish which analyses were intended before investigators had access to the relevant results. This makes later deviations visible and reduces ambiguity about whether choices were data-driven.

Prespecification should not be treated as a prohibition against improving an analysis. Researchers sometimes discover errors in a planned method, encounter unexpected data problems, or identify a better analytical approach for defensible reasons. The more informative practice is to preserve the prespecified analysis, document deviations, explain why they occurred, and distinguish planned from exploratory work.

Transparency changes the interpretation. An openly labeled sensitivity or exploratory analysis is not equivalent to quietly replacing a prespecified analysis after seeing which version produces the preferred result.

04 · A Practical Example

How One Dataset Can Generate Several Plausible Results

Hypothetical Example

Five models, one published result

Imagine a study examining whether use of an educational technology platform is associated with examination performance. The researchers have several defensible choices about covariate adjustment.

Original plan The protocol specifies adjustment for prior academic performance and year level.
Analytical exploration After seeing the data, the researchers try the unadjusted model and several models containing different combinations of demographic, academic, and engagement variables.
Results differ Most models produce small estimates with substantial uncertainty. One particular adjustment produces a larger association and p = 0.03.
Selective report Only that model appears in the article, with no indication that alternative specifications were tried or that it differed from the prespecified analysis.

The published coefficient may have been calculated perfectly. What is misleading is the evidential context. A reader sees one apparently confirmatory analysis when the researchers actually observed a range of results and selected one after examining them.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Selective Analysis Reporting

Misconception

If every analysis is statistically valid, choosing among them cannot create bias

Each analysis may be defensible in isolation while the selection process remains biased. If the observed findings determine which valid analysis is reported, readers see a result-dependent subset of the analytical possibilities.

Misconception

Researchers must never change a prespecified analysis

Legitimate changes are sometimes necessary. The relevant issues are why and when the change occurred and whether it is disclosed. Transparent deviations are fundamentally different from silently selecting an analysis because its result is more attractive.

Misconception

A p-value below 0.05 has the same meaning regardless of how many analyses were tried

No. The inferential context changes when a reported result has been selected after examining many alternatives. Conventional inferential quantities are not designed to make an undisclosed search across analytical specifications disappear.

Misconception

A protocol-publication discrepancy proves that researchers selected the favorable analysis

Not by itself. Dwan and colleagues emphasized that discrepancies can reflect bias, error, or legitimate changes. The timing, rationale, documentation, and relationship between the change and the observed results all matter.

Misconception

Sensitivity analyses create selective analysis reporting

No. Sensitivity analyses can strengthen a study by showing whether conclusions depend on reasonable analytical choices. The concern arises when researchers selectively disclose only the sensitivity analysis that supports the preferred conclusion rather than showing the relevant range of results.

06 · What This Means for You

Evaluate the Analytical Path, Not Just the Final Model

When reading a paper, you usually cannot reconstruct every analysis the researchers considered. You can, however, ask whether the reported analysis appears to have been specified before the results were known and whether important analytical alternatives are acknowledged.

A simple decision framework

If a statistical analysis plan or protocol is available
Compare the planned analytical approach with the one actually reported.
If the reported analysis differs from the original plan
Look for a dated explanation and determine whether the change could have occurred after relevant results were known.
If several analytical choices are equally plausible
Look for sensitivity analyses showing whether the substantive conclusion survives reasonable alternatives.
If an unexpected subgroup, cut-point, or adjusted model drives the conclusion
Treat its confirmatory status cautiously unless the choice was prespecified or independently replicated.

For your own work, make important analytical decisions before examining the relevant results whenever feasible. A dated statistical analysis plan can be valuable even outside fields where formal preregistration is customary. When plans must change, retain the original specification and explain the change rather than rewriting the analytical history.

Exploratory analysis remains useful. The goal is not to eliminate exploration but to identify it honestly. Readers can interpret a post hoc finding appropriately when they know that it emerged from exploration rather than from a single prespecified confirmatory test.

07 · A Quick Checklist

How to Check for Selective Analysis Reporting

When evaluating a reported analysis, check:
Whether a protocol, preregistration, or statistical analysis plan specifies the analysis in advance.
Whether the reported model matches the prespecified model and analytical population.
Whether covariates, transformations, cut-points, subgroup definitions, and missing-data methods were chosen before the relevant results were examined.
Whether deviations from the planned analysis are disclosed and justified.
Whether reasonable alternative analyses produce materially different conclusions.
Whether surprising subgroup or model-specific findings are presented as exploratory when appropriate.
Whether the paper reports enough analytical detail for readers to understand how the final result was obtained.
08 · Frequently Asked Questions

Frequently Asked Questions About Selective Analysis Reporting

What is selective analysis reporting?

It occurs when multiple analyses of the same data are possible or conducted and the analysis that becomes reported is selected based on characteristics of its result, such as statistical significance, magnitude, or direction.

Is selective analysis reporting the same as p-hacking?

They overlap but are not perfect synonyms. P-hacking usually refers to analytical or data-processing choices made in pursuit of statistically significant results. Selective analysis reporting specifically concerns which analytical results become visible after multiple possibilities exist, whether or not the researcher explicitly describes the process as searching for significance.

Are multiple analyses inherently problematic?

No. Multiple analyses may be necessary for sensitivity analysis, robustness checking, exploration, or addressing distinct questions. The concern is undisclosed result-dependent selection rather than multiplicity itself.

Does preregistration eliminate selective analysis reporting?

It can make selective changes easier to detect, but registration alone does not guarantee adherence or complete reporting. The prespecified record must contain enough analytical detail to make meaningful comparison possible, and deviations still need to be reported transparently.

How can I tell whether an analysis was selected after seeing the data?

Often you cannot determine this from the publication alone. Compare the paper with dated protocols, registrations, statistical analysis plans, or other records created before the relevant results were available.

Can selective analysis reporting affect a meta-analysis?

Yes. If the effect estimate extracted from each study has itself been selected from multiple alternatives based on its result, a meta-analysis can combine already-selected estimates and inherit that bias.

09 · The Bottom Line

The Final Analysis Does Not Always Reveal the Analytical Journey

The Bottom Line

Selective analysis reporting can distort research when the result presented to readers is chosen from multiple analytical possibilities because it looks more favorable, significant, or compelling.

The existence of several possible analyses is not itself the problem. What matters is whether analytical decisions were driven by the findings and hidden from readers. Prespecification, transparent deviations, and appropriately reported sensitivity analyses make that distinction much easier to evaluate.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes