Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Judge a Statistical Result You Cannot Reproduce?

Being unable to reproduce a statistical result does not automatically make it wrong. Judge why reproduction failed, what evidence remains independently assessable, and how much your conclusion depends on the unverified analysis.

384
Judging a Result You Cannot Reproduce Guide 384 of 899
01 · The Question

What Can You Conclude When You Cannot Reproduce the Result?

You have examined a paper carefully, perhaps even tried to reconstruct one of its analyses, but you cannot obtain the reported result. Maybe the raw data are unavailable. Perhaps the authors did not share their code, or the statistical model is described without enough detail to rebuild it. In other cases, you have the necessary materials but your calculation simply produces a different answer.

That creates an uncomfortable appraisal problem. Should you distrust the result, accept it provisionally, or conclude that the analysis is wrong?

The answer depends heavily on why reproduction failed. Failure caused by unavailable data is not equivalent to failure caused by executable code that produces different results. Critical appraisal therefore requires you to identify what is missing or inconsistent before deciding how much the problem should affect your confidence in the finding.

02 · The Short Answer

Non-Reproduction Is a Reason for Scrutiny, Not an Automatic Refutation

In Brief

If you cannot reproduce a statistical result, first determine why: insufficient information, unavailable data or code, ambiguous analytical choices, technical problems, or a genuine discrepancy can have very different implications. Then judge the result using what remains independently assessable, including the study design, reported methods, effect estimate, uncertainty, internal consistency, sensitivity analyses, and transparency of the analytical process.

An unexplained failure to reproduce an important result should reduce how confidently you rely on it, especially when the central claim depends heavily on that analysis. But inability to reproduce a result from the published paper alone is not evidence by itself that the result is false.

03 · What You Need to Know

Diagnose the Reproducibility Problem Before Judging the Finding

Start by defining what you tried to reproduce

Reproducibility is often discussed as though it were a single yes-or-no property. In practice, your appraisal should be more precise.

You might be trying to verify a percentage reported in a table, reconstruct an effect estimate from summary data, rerun a regression using shared data and code, or reproduce every table and figure in the article. Those are very different tasks.

For computational results, reproduction generally means obtaining the same or sufficiently similar results by applying the reported analysis to the same underlying data. This is distinct from replication, which asks whether a scientific claim can be supported with new data. The distinction matters because failure at one does not necessarily establish failure at the other.

Reproduction problem You cannot obtain the reported result from the original data, code, methods, or information available to you.
Replication problem A new investigation does not produce evidence consistent with the original scientific claim.

There are several fundamentally different reasons reproduction can fail

Before interpreting the failure, locate the obstacle.

What happened? What it tells you What it does not establish
Raw data are unavailable Independent computational verification is restricted That the reported result is incorrect
Code is unavailable Some analytical decisions may be difficult to reconstruct That the authors used inappropriate code
Methods are insufficiently specified Reporting or analytical transparency may be inadequate for reproduction Which particular analytical decision, if any, is wrong
Available code will not execute The computational workflow may have technical dependencies or documentation problems That the statistical result itself is necessarily incorrect
Code executes but produces different results There is a substantive discrepancy requiring explanation Which version is correct without further investigation
Minor numerical differences occur Rounding, software versions, numerical precision, or implementation details may matter That the scientific conclusion necessarily changes

These scenarios should not receive the same appraisal judgment. A paper that cannot be reproduced because confidential patient-level data cannot legally be released presents a different problem from a paper whose publicly archived code and data generate estimates inconsistent with its published tables.

Ask whether the result is reproducible in principle

Some barriers to reproduction are legitimate. Individual-level health, educational, administrative, commercial, or otherwise sensitive data may be subject to ethical, legal, contractual, or privacy restrictions. In such circumstances, unrestricted public release may be inappropriate.

The relevant appraisal question becomes whether the authors provide enough transparency to understand what was done and whether appropriate mechanisms exist for qualified access or verification where feasible.

Conversely, when an analysis depends on custom computational procedures but neither code nor sufficient methodological detail is available, uncertainty about the analytical workflow increases. Code can contain consequential decisions that are difficult to communicate completely in prose. Making data, code, and outputs explicitly connected can make computational verification substantially easier.

Examine what you can verify without reproducing the computation

A failed reproduction attempt does not eliminate ordinary critical appraisal. You can still inspect whether the design addresses the research question, whether the statistical method appears appropriate, whether the variables and outcomes are clearly defined, whether the sample and exclusions are reported, and whether the analysis appears consistent with the stated design.

Then inspect the numerical evidence. Are denominators consistent? Do descriptive statistics agree across tables? Does the direction of the reported effect match the raw summaries? Are estimates accompanied by appropriate measures of uncertainty? Do the conclusions accurately represent the results?

This is one reason you generally do not need to recalculate every statistic during critical appraisal. Statistical reproduction can strengthen verification, but substantial appraisal remains possible without it.

Judge the effect estimate and its uncertainty, not just whether the P-value can be recreated

A reported result should not be reduced to whether its P-value falls above or below 0.05. Cochrane guidance emphasizes interpreting the point estimate together with its confidence interval because the interval conveys information about statistical uncertainty and precision.

That remains useful even when you cannot reproduce the underlying calculation. Ask whether the reported estimate represents a trivial, moderate, or potentially important effect in the substantive context. Examine how wide the confidence interval is and whether it includes materially different interpretations.

The relationship among P-values, effect sizes, and confidence intervals can also expose internal inconsistencies. For corresponding conventional tests and intervals, certain combinations should agree logically. Apparent contradictions deserve investigation, although differences in statistical procedures can sometimes explain them.

Look for robustness rather than one privileged analysis

A central estimate becomes more persuasive when reasonable analytical alternatives produce substantively similar conclusions. Sensitivity analyses may vary assumptions, definitions, exclusions, missing-data procedures, model specifications, or other defensible analytical decisions.

This does not mean that every possible analysis should produce the same number. Statistical analyses often contain legitimate researcher choices, and different reasonable specifications may yield different estimates. What matters is whether the substantive conclusion is unusually dependent on one narrow set of decisions.

If a claim appears only under a particular specification while plausible alternatives produce substantially different conclusions, that fragility is relevant even if the preferred model was computed correctly.

Consider whether the paper gives you enough information to understand the analysis

Statistical reporting guidelines such as SAMPL emphasize reporting statistical methods and analyses sufficiently clearly. Transparency is not merely an editorial nicety. Without adequate information about how variables were treated, models were specified, missing observations were handled, or adjustments were made, readers may be unable to determine what generated the published result.

When the analysis is sophisticated, the absence of sufficient information becomes particularly consequential. If you cannot determine whether the method was applied appropriately, it may be better to recognize the limits of your own evaluation rather than infer validity from technical complexity.

The importance of non-reproduction depends on how much rests on the result

Not every discrepancy deserves the same response. A rounding difference in a secondary descriptive table is unlikely to overturn an otherwise coherent study. Failure to reproduce the primary analysis on which the paper's main conclusion depends is much more consequential.

Ask a counterfactual question: if this result were substantially different, would your interpretation of the study change?

If yes, unresolved reproducibility becomes central to the appraisal. If no, document the problem but keep its importance proportional to its role in the study.

Watch Out

Do not convert “I could not reproduce this result” into “this result has been proven wrong.” State what you attempted, what information or materials you used, where the attempt failed, and whether the discrepancy is large enough to affect the scientific interpretation.

Reproducibility itself is not the same as scientific validity

A result can be perfectly reproducible and still arise from a biased design, inappropriate model, poorly measured outcome, unjustified adjustment strategy, or selective analysis. Reproduction asks whether the computation can be obtained again. Critical appraisal asks the broader question of whether the evidence justifies the inference.

The reverse also deserves caution. A result that cannot currently be reproduced may ultimately prove correct once missing code, data transformations, software dependencies, or analytical details are supplied.

04 · A Practical Example

What If Shared Data Produce a Different Result?

Hypothetical Example

A regression coefficient does not match the published table

Suppose a paper reports an adjusted association of 0.42 with a 95% confidence interval of 0.18 to 0.66. The authors provide a dataset and partial analysis code. You run the available code and obtain an estimate of 0.29 with a confidence interval of 0.04 to 0.54.

Identify the discrepancy. The result is not merely different in the final decimal place. Both the point estimate and confidence interval differ enough to warrant investigation.
Check whether you actually reproduced the same analysis. Compare the analytic sample, exclusions, variable coding, transformations, reference categories, missing-data procedures, covariates, weighting, clustering, and variance estimator.
Inspect the available documentation. Determine whether the paper, supplement, repository, or code explains the discrepancy. Partial code may omit preprocessing or model-building steps that generated the published analysis.
Ask whether the scientific interpretation changes. In this hypothetical example, both estimates remain positive, but their magnitudes differ. Whether that matters depends on what size of association would be substantively important and how the study's claims were framed.
Report the uncertainty accurately. If no explanation can be found, the defensible conclusion is that you could not reproduce the published estimate from the supplied materials. You have identified an unresolved discrepancy, not automatically established which estimate is correct.

If the difference altered the direction of the association, moved the estimate from substantively important to trivial, or materially changed the uncertainty around it, the unresolved discrepancy would deserve considerably more weight in your appraisal.

05 · What Researchers Often Get Wrong

Common Mistakes When a Result Cannot Be Reproduced

Misconception

Failure to Reproduce Means the Original Result Is False

Not necessarily. You may lack data, code, preprocessing instructions, model details, software dependencies, or other information needed to reproduce the analysis. A genuine numerical contradiction is more concerning than an attempt that could never fully reconstruct the original workflow.

Misconception

If the Data Are Available, Reproduction Should Be Automatic

Data alone may not be enough. Analytical results can depend on data cleaning, recoding, exclusions, transformations, model specifications, software procedures, and other decisions. Code and documentation can therefore be as important as the dataset for computational reproduction.

Misconception

A Reproducible Result Must Be a Valid Result

Reproducibility establishes something about the computational pathway, not the entire evidential chain. The same code can repeatedly reproduce an estimate from a biased study or an inappropriate statistical model.

Misconception

Any Numerical Difference Is Scientifically Important

Minor differences may arise from rounding, software versions, numerical precision, or implementation details without changing the interpretation. Focus on whether the discrepancy alters the magnitude, direction, precision, or substantive conclusion.

Misconception

You Should Trust the Published Number Until Someone Proves It Wrong

Publication is relevant context, but it does not eliminate uncertainty. If an important result cannot be independently checked because essential analytical information is missing, that limitation belongs in your appraisal. The appropriate response is calibrated uncertainty rather than automatic acceptance or rejection.

06 · What This Means for You

Match Your Confidence to What Can Actually Be Verified

When reproduction fails, resist the temptation to make the problem binary. Instead, identify where the evidential chain becomes uncertain and ask how important that uncertainty is to your intended use of the paper.

A simple decision framework

If reproduction is impossible because the required data or code are unavailable
Appraise everything that remains visible, document the transparency limitation, and avoid treating non-reproduction itself as evidence that the result is wrong.
If methods are ambiguous but the result appears internally coherent
Treat the uncertainty about analytical choices as a limitation and look for sensitivity analyses or additional documentation.
If the supplied data and code produce a materially different result
Investigate preprocessing, model specification, software dependencies, and other plausible explanations. If the discrepancy remains unexplained, reduce reliance on the affected finding accordingly.
If the disputed analysis is central, technically complex, or consequential for an important decision

The central principle is proportionality. A result that cannot be reproduced deserves scrutiny, but the strength of your concern should reflect why reproduction failed, how large the discrepancy is, and how much of the paper's conclusion depends on that result.

07 · A Quick Checklist

When You Cannot Reproduce a Statistical Result

Before deciding how much to trust the result, check:
What exactly were you trying to reproduce: a simple statistic, one model, or the complete analysis?
Do you have the same underlying data or only summary information from the article?
Are the analysis code, preprocessing steps, exclusions, variable definitions, and model specifications available?
Could rounding, software versions, missing-data handling, weighting, transformations, or different analytic samples explain the discrepancy?
Are the published estimates, confidence intervals, P-values, tables, and narrative descriptions internally coherent?
Do sensitivity or robustness analyses support a similar substantive conclusion under reasonable alternatives?
Would the discrepancy materially change the magnitude, direction, precision, or practical interpretation of the result?
Is the result central enough to your decision that independent statistical expertise is warranted?
08 · Frequently Asked Questions

Questions About Results You Cannot Reproduce

Does non-reproducibility mean research misconduct occurred?

No. Reproduction can fail for many reasons, including incomplete documentation, unavailable data, software dependencies, coding errors, differing analytical decisions, or genuine mistakes. Evidence of non-reproduction alone does not establish misconduct or intent.

Can I judge a statistical result without the raw data?

Yes, although your appraisal will have limits. You can examine study design, reported methods, descriptive statistics, effect estimates, uncertainty, internal consistency, sensitivity analyses, and the relationship between the results and conclusions. Raw data would permit additional checks that the paper alone cannot support.

Is missing analysis code a red flag?

It can limit verification, particularly when results depend heavily on custom or complex computational procedures, but its meaning depends on disciplinary practices, journal policies, legitimate restrictions, and how completely the methods are otherwise documented. Treat it as a transparency and reproducibility consideration rather than automatic evidence of an incorrect analysis.

What if I obtain almost the same result?

Ask whether the difference is scientifically consequential. Small numerical discrepancies may arise from rounding, numerical precision, software versions, or minor implementation differences. Exact numerical identity is less important than determining whether the discrepancy changes the substantive interpretation.

What if the authors' code produces a different result from the paper?

That deserves investigation. Confirm that you used the intended data, code version, software environment, and execution sequence. If the discrepancy persists and materially affects the reported finding, it becomes a substantive limitation of the result until adequately explained.

Should I contact the authors?

For a consequential unresolved discrepancy, contacting the corresponding author can be reasonable. Describe exactly what you attempted, which materials and software you used, the result you obtained, and where it differs from the publication rather than beginning with an accusation of error.

Is reproducibility more important than robustness?

They answer different questions. Reproducibility asks whether the reported computation can be obtained again, whereas robustness asks whether the substantive conclusion persists under reasonable alternative analytical decisions. A strong appraisal may need to consider both.

09 · The Bottom Line

Unexplained Non-Reproduction Should Change Your Certainty, Not End the Appraisal

The Bottom Line

When you cannot reproduce a statistical result, determine why before deciding what the failure means. Missing materials, incomplete reporting, technical obstacles, and a genuine numerical contradiction are different problems and should not be treated as equivalent evidence against the result.

Judge what remains independently assessable and ask whether the unresolved problem could materially change the study's conclusion. The more consequential the result and the more direct the unexplained discrepancy, the more cautiously you should rely on it.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes