01 · The Question
What Can You Conclude When You Cannot Reproduce the Result?
You have examined a paper carefully, perhaps even tried to reconstruct one of its analyses, but you cannot obtain the reported result. Maybe the raw data are unavailable. Perhaps the authors did not share their code, or the statistical model is described without enough detail to rebuild it. In other cases, you have the necessary materials but your calculation simply produces a different answer.
That creates an uncomfortable appraisal problem. Should you distrust the result, accept it provisionally, or conclude that the analysis is wrong?
The answer depends heavily on why reproduction failed. Failure caused by unavailable data is not equivalent to failure caused by executable code that produces different results. Critical appraisal therefore requires you to identify what is missing or inconsistent before deciding how much the problem should affect your confidence in the finding.
03 · What You Need to Know
Diagnose the Reproducibility Problem Before Judging the Finding
Start by defining what you tried to reproduce
Reproducibility is often discussed as though it were a single yes-or-no property. In practice, your appraisal should be more precise.
You might be trying to verify a percentage reported in a table, reconstruct an effect estimate from summary data, rerun a regression using shared data and code, or reproduce every table and figure in the article. Those are very different tasks.
For computational results, reproduction generally means obtaining the same or sufficiently similar results by applying the reported analysis to the same underlying data. This is distinct from replication, which asks whether a scientific claim can be supported with new data. The distinction matters because failure at one does not necessarily establish failure at the other.
Reproduction problem
You cannot obtain the reported result from the original data, code, methods, or information available to you.
Replication problem
A new investigation does not produce evidence consistent with the original scientific claim.
There are several fundamentally different reasons reproduction can fail
Before interpreting the failure, locate the obstacle.
| What happened? |
What it tells you |
What it does not establish |
| Raw data are unavailable |
Independent computational verification is restricted |
That the reported result is incorrect |
| Code is unavailable |
Some analytical decisions may be difficult to reconstruct |
That the authors used inappropriate code |
| Methods are insufficiently specified |
Reporting or analytical transparency may be inadequate for reproduction |
Which particular analytical decision, if any, is wrong |
| Available code will not execute |
The computational workflow may have technical dependencies or documentation problems |
That the statistical result itself is necessarily incorrect |
| Code executes but produces different results |
There is a substantive discrepancy requiring explanation |
Which version is correct without further investigation |
| Minor numerical differences occur |
Rounding, software versions, numerical precision, or implementation details may matter |
That the scientific conclusion necessarily changes |
These scenarios should not receive the same appraisal judgment. A paper that cannot be reproduced because confidential patient-level data cannot legally be released presents a different problem from a paper whose publicly archived code and data generate estimates inconsistent with its published tables.
Ask whether the result is reproducible in principle
Some barriers to reproduction are legitimate. Individual-level health, educational, administrative, commercial, or otherwise sensitive data may be subject to ethical, legal, contractual, or privacy restrictions. In such circumstances, unrestricted public release may be inappropriate.
The relevant appraisal question becomes whether the authors provide enough transparency to understand what was done and whether appropriate mechanisms exist for qualified access or verification where feasible.
Conversely, when an analysis depends on custom computational procedures but neither code nor sufficient methodological detail is available, uncertainty about the analytical workflow increases. Code can contain consequential decisions that are difficult to communicate completely in prose. Making data, code, and outputs explicitly connected can make computational verification substantially easier.
Examine what you can verify without reproducing the computation
A failed reproduction attempt does not eliminate ordinary critical appraisal. You can still inspect whether the design addresses the research question, whether the statistical method appears appropriate, whether the variables and outcomes are clearly defined, whether the sample and exclusions are reported, and whether the analysis appears consistent with the stated design.
Then inspect the numerical evidence. Are denominators consistent? Do descriptive statistics agree across tables? Does the direction of the reported effect match the raw summaries? Are estimates accompanied by appropriate measures of uncertainty? Do the conclusions accurately represent the results?
This is one reason you generally do not need to recalculate every statistic during critical appraisal. Statistical reproduction can strengthen verification, but substantial appraisal remains possible without it.
Judge the effect estimate and its uncertainty, not just whether the P-value can be recreated
A reported result should not be reduced to whether its P-value falls above or below 0.05. Cochrane guidance emphasizes interpreting the point estimate together with its confidence interval because the interval conveys information about statistical uncertainty and precision.
That remains useful even when you cannot reproduce the underlying calculation. Ask whether the reported estimate represents a trivial, moderate, or potentially important effect in the substantive context. Examine how wide the confidence interval is and whether it includes materially different interpretations.
The relationship among P-values, effect sizes, and confidence intervals can also expose internal inconsistencies. For corresponding conventional tests and intervals, certain combinations should agree logically. Apparent contradictions deserve investigation, although differences in statistical procedures can sometimes explain them.
Look for robustness rather than one privileged analysis
A central estimate becomes more persuasive when reasonable analytical alternatives produce substantively similar conclusions. Sensitivity analyses may vary assumptions, definitions, exclusions, missing-data procedures, model specifications, or other defensible analytical decisions.
This does not mean that every possible analysis should produce the same number. Statistical analyses often contain legitimate researcher choices, and different reasonable specifications may yield different estimates. What matters is whether the substantive conclusion is unusually dependent on one narrow set of decisions.
If a claim appears only under a particular specification while plausible alternatives produce substantially different conclusions, that fragility is relevant even if the preferred model was computed correctly.
Consider whether the paper gives you enough information to understand the analysis
Statistical reporting guidelines such as SAMPL emphasize reporting statistical methods and analyses sufficiently clearly. Transparency is not merely an editorial nicety. Without adequate information about how variables were treated, models were specified, missing observations were handled, or adjustments were made, readers may be unable to determine what generated the published result.
When the analysis is sophisticated, the absence of sufficient information becomes particularly consequential. If you cannot determine whether the method was applied appropriately, it may be better to recognize the limits of your own evaluation rather than infer validity from technical complexity.
The importance of non-reproduction depends on how much rests on the result
Not every discrepancy deserves the same response. A rounding difference in a secondary descriptive table is unlikely to overturn an otherwise coherent study. Failure to reproduce the primary analysis on which the paper's main conclusion depends is much more consequential.
Ask a counterfactual question: if this result were substantially different, would your interpretation of the study change?
If yes, unresolved reproducibility becomes central to the appraisal. If no, document the problem but keep its importance proportional to its role in the study.
Watch Out
Do not convert “I could not reproduce this result” into “this result has been proven wrong.” State what you attempted, what information or materials you used, where the attempt failed, and whether the discrepancy is large enough to affect the scientific interpretation.
Reproducibility itself is not the same as scientific validity
A result can be perfectly reproducible and still arise from a biased design, inappropriate model, poorly measured outcome, unjustified adjustment strategy, or selective analysis. Reproduction asks whether the computation can be obtained again. Critical appraisal asks the broader question of whether the evidence justifies the inference.
The reverse also deserves caution. A result that cannot currently be reproduced may ultimately prove correct once missing code, data transformations, software dependencies, or analytical details are supplied.
04 · A Practical Example
What If Shared Data Produce a Different Result?
Hypothetical Example
A regression coefficient does not match the published table
Suppose a paper reports an adjusted association of 0.42 with a 95% confidence interval of 0.18 to 0.66. The authors provide a dataset and partial analysis code. You run the available code and obtain an estimate of 0.29 with a confidence interval of 0.04 to 0.54.
Identify the discrepancy. The result is not merely different in the final decimal place. Both the point estimate and confidence interval differ enough to warrant investigation.
Check whether you actually reproduced the same analysis. Compare the analytic sample, exclusions, variable coding, transformations, reference categories, missing-data procedures, covariates, weighting, clustering, and variance estimator.
Inspect the available documentation. Determine whether the paper, supplement, repository, or code explains the discrepancy. Partial code may omit preprocessing or model-building steps that generated the published analysis.
Ask whether the scientific interpretation changes. In this hypothetical example, both estimates remain positive, but their magnitudes differ. Whether that matters depends on what size of association would be substantively important and how the study's claims were framed.
Report the uncertainty accurately. If no explanation can be found, the defensible conclusion is that you could not reproduce the published estimate from the supplied materials. You have identified an unresolved discrepancy, not automatically established which estimate is correct.
If the difference altered the direction of the association, moved the estimate from substantively important to trivial, or materially changed the uncertainty around it, the unresolved discrepancy would deserve considerably more weight in your appraisal.
07 · A Quick Checklist
When You Cannot Reproduce a Statistical Result
Before deciding how much to trust the result, check:
What exactly were you trying to reproduce: a simple statistic, one model, or the complete analysis?
Do you have the same underlying data or only summary information from the article?
Are the analysis code, preprocessing steps, exclusions, variable definitions, and model specifications available?
Could rounding, software versions, missing-data handling, weighting, transformations, or different analytic samples explain the discrepancy?
Are the published estimates, confidence intervals, P-values, tables, and narrative descriptions internally coherent?
Do sensitivity or robustness analyses support a similar substantive conclusion under reasonable alternatives?
Would the discrepancy materially change the magnitude, direction, precision, or practical interpretation of the result?
Is the result central enough to your decision that independent statistical expertise is warranted?