01 · The Question
Must You Reproduce the Numbers Before You Can Trust Them?
You are critically appraising a paper and reach the Results section. There are regression coefficients, confidence intervals, P-values, adjusted models, perhaps several tables of secondary analyses. Do you need to open statistical software and reproduce the calculations before you can judge the study?
Usually, no. Critical appraisal and statistical reproduction are related activities, but they are not the same thing. A careful reader can identify many consequential statistical strengths and weaknesses without access to the raw data or the ability to rerun the analysis.
The more useful question is not whether you can reproduce every number. It is whether the evidence reported in the paper is sufficient to judge whether the analysis was appropriate, internally coherent, sufficiently precise, and interpreted in a way that the results actually support.
03 · What You Need to Know
Critical Appraisal Is Not the Same as Statistical Reproduction
Start by asking whether the analysis answers the research question
A calculation can be mathematically correct and still be a poor analysis. If the wrong statistical model was chosen, an important source of bias was ignored, observations that should not be treated as independent were analyzed as though they were, or an inappropriate outcome was tested, reproducing the arithmetic merely reproduces the underlying problem.
This is why critical appraisal begins upstream of the final numbers. Ask what was estimated, why that quantity was appropriate for the research question, how the data arose, what assumptions the method requires, and whether the study appears to satisfy those assumptions.
You can often evaluate much of this from the Methods and Results sections. For example, you might ask whether a statistical test fits the type of outcome, whether repeated observations or clustering were handled appropriately, whether adjusted analyses used defensible covariates, and whether the reported analyses correspond to the questions the researchers claimed they were investigating.
Separate checking, recalculation, reproduction, and reanalysis
These activities are easy to blur together, but they demand very different amounts of information and statistical work.
| Activity |
What you are doing |
What you may need |
| Consistency checking |
Looking for contradictions among numbers, tables, figures, text, and conclusions |
Usually the published paper |
| Selective recalculation |
Recomputing a particular statistic from reported information |
Sufficient summary statistics and knowledge of the calculation |
| Statistical reproduction |
Trying to obtain the reported analysis using the authors' data and procedures |
Usually the analytic data, model specification, and enough procedural detail |
| Reanalysis |
Analyzing the data again, potentially using different decisions or methods |
Data, substantial methodological information, and an explicit analytical rationale |
A reader performing critical appraisal will usually spend most of the time on the first activity and on conceptual evaluation of the methods. The others become useful when the appraisal question requires them.
Many statistical problems are visible without recalculation
You can learn a surprising amount by examining the reported results carefully. Check whether sample sizes agree across the abstract, tables, and Results section. Look at whether percentages correspond plausibly to the stated denominators. Ask whether the direction of an effect estimate agrees with the prose describing it. Compare estimates with their confidence intervals and inspect whether the reported P-values appear compatible with them.
For many commonly reported effect measures, the null value is straightforward. Difference measures such as a mean difference or risk difference generally have a null value of 0, while ratio measures such as risk ratios and odds ratios generally have a null value of 1. Effect estimates should also be accompanied by information about uncertainty, commonly a confidence interval. Cochrane guidance emphasizes interpreting point estimates together with their confidence intervals rather than relying on a binary significant versus nonsignificant distinction.
These relationships let you perform useful plausibility checks before touching a calculator. For a conventional two-sided test, for example, a 95% confidence interval for an effect that excludes its null value generally corresponds to a P-value below 0.05 when the interval and hypothesis test are based on corresponding methods. That relationship has qualifications, so it should be used as a consistency check rather than as a universal law for every statistical procedure.
Do not reduce appraisal to whether P < 0.05
Even a perfectly reproduced P-value does not tell you whether the finding is important. P-values do not directly quantify the magnitude of an effect, and statistical significance should not be treated as a substitute for substantive importance. Confidence intervals add information about the range and precision of the estimated effect.
When appraising a result, therefore, examine the estimate itself, its direction and magnitude, and the uncertainty surrounding it. This is why considering P-values alongside effect sizes and confidence intervals is generally more informative than verifying whether the authors crossed an arbitrary significance threshold.
This distinction becomes particularly important when a statistically significant finding may nevertheless be substantively unimportant, or when a nonsignificant result still contains useful information about plausible effects.
Sometimes a quick recalculation is genuinely useful
Recalculation becomes more valuable when the necessary information is available and something specific needs checking. Simple examples include recomputing a percentage from a numerator and denominator, checking a reported difference between two means, verifying an unadjusted risk ratio from a two-by-two table, or examining whether a reported confidence interval and P-value appear mutually compatible.
Published summary information can sometimes support additional calculations. Cochrane, for example, documents methods for deriving standard errors from confidence intervals or exact P-values for certain effect estimates, while Altman and Bland describe circumstances in which P-values and confidence intervals can be approximately derived from one another. Such procedures depend on the statistic, assumptions, and information reported, so they should not be applied mechanically.
A failed recalculation does not automatically prove the paper is wrong
Suppose you try to reconstruct a statistic and obtain a different answer. That discrepancy deserves investigation, but several explanations are possible. The authors may have used adjusted rather than crude estimates, a transformation, a different variance estimator, a particular treatment of missing data, weighting, clustering, rounding, or a model specification that cannot be reconstructed from the table alone.
The appropriate conclusion is initially modest: the reported result could not be reproduced from the information available to you. That is different from demonstrating that the result is erroneous. The next question is how much confidence to place in a result you cannot reproduce and whether the unexplained discrepancy matters to your use of the study.
Some analyses cannot realistically be verified from the article alone
Modern analyses may involve multivariable regression, multilevel models, multiple imputation, complex survey weights, penalized prediction models, time-to-event methods, Bayesian models, or extensive preprocessing. A journal article may describe these methods adequately for interpretation while still not contain everything required to reproduce every numerical result.
In those situations, trying to reverse-engineer the entire analysis from a results table can create false confidence. You may be assuming analytical choices that the authors did not actually make.
Instead, determine what you can evaluate confidently. You can inspect the design, variables, outcome definition, stated model, sample size, missing-data procedures, reported estimates, uncertainty, sensitivity analyses, and consistency between the results and conclusions. When the method itself exceeds your statistical competence, recognizing the limits of what you can evaluate confidently is part of rigorous appraisal, not a failure of it.
Recalculation cannot repair missing information
A calculator cannot reveal analyses that were never reported. If a paper presents many outcomes, models, subgroups, or statistical tests, the important concern may be the analytical process rather than whether each displayed P-value is numerically correct.
For example, a paper reporting many statistical tests raises questions about multiplicity and how the analyses were planned and interpreted. Similarly, recalculating the results that appear in print cannot by itself determine whether nonsignificant analyses were omitted. That requires attention to protocols, registrations, analysis plans, and patterns of reporting.
Watch Out
A successful recalculation verifies a calculation under the assumptions and inputs you used. It does not establish that the study design was unbiased, the model was appropriate, all analyses were reported, the data were accurate, or the authors' interpretation was justified.
04 · A Practical Example
How Far Should You Go When a Result Looks Suspicious?
Hypothetical Example
A trial reports an apparently important treatment difference
Imagine that a paper compares an intervention with a control condition. The primary outcome is reported as 52 events among 400 participants in the intervention group and 68 events among 400 participants in the control group. The authors provide an effect estimate, confidence interval, and P-value and describe the result as evidence of a meaningful benefit.
First: inspect the research design. Before calculating anything, ask whether allocation, follow-up, outcome measurement, exclusions, and missing data could introduce important bias. A correct risk ratio cannot rescue a seriously biased comparison.
Second: inspect the reported quantities. Check whether the group totals and event counts agree across the abstract, text, tables, and participant flow. Determine which effect measure the authors actually report and whether its direction matches the raw event rates.
Third: evaluate magnitude and uncertainty. Examine the absolute event rates, reported effect estimate, and confidence interval. Ask whether the range of effects compatible with the data includes effects that would lead to materially different interpretations.
Fourth: recalculate only if it answers a useful question. Because the event counts and denominators are available, you could independently calculate simple unadjusted quantities such as the two event proportions and their crude difference or ratio. This may help detect transcription or reporting inconsistencies.
Finally: distinguish verification from appraisal. Even if your crude calculation agrees with the paper, you still need to judge bias, precision, analytical choices, multiplicity where relevant, and whether the authors' conclusion is proportionate to the evidence.
The recalculation is useful because it answers a targeted question: are simple reported quantities consistent with the underlying numbers? It is not a ritual that must be performed before any statistical result can be discussed.
07 · A Quick Checklist
Before You Start Recalculating a Paper's Statistics
Before recalculating, check:
Can you identify exactly what statistical quantity the authors estimated?
Is the statistical method appropriate for the research question, study design, and type of data?
Do sample sizes, percentages, estimates, tables, figures, and textual descriptions appear internally consistent?
Have you examined the magnitude and direction of the effect rather than focusing only on statistical significance?
Have you examined the confidence interval or another measure of uncertainty around the estimate?
Do you have enough reported information to perform the recalculation legitimately?
If your result differs, have you checked whether adjustment, transformations, weighting, missing-data handling, or rounding could explain the discrepancy?
Would resolving the numerical discrepancy materially change how you interpret or use the study?
If the method is beyond your expertise and the conclusion matters, have you considered obtaining specialist statistical input?