03 · What You Need to Know
Statistical Consultation Should Resolve a Consequential Uncertainty
Start by identifying exactly what you cannot judge
"The statistics are complicated" is not yet a useful consultation question.
Try to identify the unresolved issue precisely. Perhaps you cannot determine whether the authors adjusted for an inappropriate post-exposure variable. Maybe a multilevel model appears sensible but you cannot evaluate its specification. Perhaps the paper reports excellent prediction performance but you are unsure whether the validation procedure allowed information leakage. Or the result depends heavily on a missing-data assumption whose consequences you cannot assess.
This distinction matters because statistical expertise is most useful when directed toward a defined methodological problem. The American Statistical Association describes statistical consultation as obtaining professional help in selecting and using methods for obtaining and analyzing data, and its ethical guidance emphasizes competence, transparency, valid interpretation, and appropriate collaboration in statistical practice.
Do not seek consultation merely because a method has a frightening name
Unfamiliarity alone is not a reason to distrust an analysis or automatically escalate it to a specialist.
You can often evaluate much of a paper without mastering every technical detail. Study design, sampling, outcome measurement, missing data, attrition, prespecification, effect magnitude, uncertainty, internal consistency, selective reporting, and the relationship between results and conclusions may remain accessible.
If you encounter a method outside your experience, first apply the principles for evaluating an advanced statistical method you cannot assess completely. Consultation becomes more valuable when the unresolved technical component remains important after that appraisal.
Ask whether the unresolved issue is central to the paper's conclusion
Statistical uncertainty deserves different levels of attention depending on where it sits in the evidential chain.
Suppose a complicated exploratory analysis appears in a supplementary appendix but does not affect the paper's central conclusion. You may be able to document that you cannot evaluate it fully and continue.
Now suppose the entire conclusion depends on one causal model whose assumptions you cannot assess. That is different. If the model is inappropriate, the main claim may change substantially.
| Situation |
Need for specialist input |
| Unfamiliar method used only for a peripheral exploratory analysis |
May be unnecessary if it does not affect your use of the study |
| Primary conclusion depends on a complex model you cannot evaluate |
Stronger reason to seek relevant expertise |
| Simple analysis with transparent methods and coherent results |
Specialist review may add little for ordinary appraisal |
| Major unexplained statistical inconsistencies |
Consultation may help determine whether they are consequential |
| Study will materially influence a high-consequence decision |
A lower threshold for independent specialist review is reasonable |
Seek help when important results are unusually sensitive to analytical choices
Suppose an association is substantial in one model, nearly disappears in another, and reverses direction after a third adjustment. Or a result is statistically significant only under one treatment of missing data, one variable transformation, or one subgroup definition.
Such sensitivity does not automatically mean the analysis is wrong. Different models may legitimately estimate different quantities or rely on different assumptions.
But when your interpretation depends on deciding which specification is defensible, specialist review may be valuable. The question is no longer merely how to read a coefficient. It concerns the relationship between modelling choices and the scientific inference.
Complex statistical adjustment may warrant specialist causal expertise
Observational causal analyses can involve propensity scores, inverse-probability weighting, matching, instrumental variables, marginal structural models, mediation analyses, target-trial approaches, or other specialized methods.
The mathematics is only part of the challenge. Valid causal interpretation may depend on assumptions about confounding, positivity, consistency, temporal ordering, model specification, selection, measurement, and the causal structure connecting variables.
If the central conclusion depends on adjustment choices you cannot evaluate, revisit whether the statistical adjustment matches the intended causal question. When the remaining uncertainty concerns specialized causal assumptions or estimation, expertise in causal inference may be more useful than generic statistical familiarity.
Prediction models can require a different kind of statistical review
A prediction-model paper may use familiar-looking regression or sophisticated machine learning, but the important appraisal issues include model complexity, sample size, data leakage, tuning, internal validation, optimism, calibration, and external performance.
If a model will be used to classify individuals, estimate risk, allocate resources, or support consequential decisions, apparently excellent accuracy should not be accepted without understanding how it was obtained.
When you cannot determine whether development and validation were methodologically sound, expertise in prediction modelling may be warranted. The warning signs for overfitting in prediction-model research can help identify where specialist scrutiny should focus.
Unexplained contradictions between reported statistics can justify escalation
Suppose confidence intervals appear incompatible with reported P-values, denominators vary without explanation, adjusted estimates cannot be reconciled with the stated model, or shared code and data produce materially different results.
Begin with ordinary consistency checking and targeted recalculation where feasible. You generally do not need to recalculate every statistic to appraise a paper.
If a central result still cannot be reconciled, specialist review can help determine whether the discrepancy arises from a legitimate analytical detail, incomplete reporting, computational error, or a more serious methodological problem. When a result cannot be reproduced from the available information, the reason for that failure matters greatly.
Missing data can become a specialist issue when conclusions depend on unverifiable assumptions
Missing observations are common, and not every incomplete dataset requires a statistician. Concern rises when a substantial amount of outcome or predictor information is missing, missingness differs importantly between groups, sophisticated imputation or weighting methods are used, or conclusions change across assumptions about the missing values.
Some assumptions about missing data cannot be verified directly from the observed dataset. Sensitivity analyses may therefore be central to judging robustness.
If the paper's main conclusion depends on technical missing-data methods whose assumptions you cannot evaluate, statistical expertise can materially improve the appraisal.
Multiplicity and selective analysis can require more than checking individual P-values
When a study contains numerous outcomes, models, time points, subgroup analyses, or interim looks at the data, evaluating individual P-values may miss the larger inferential problem.
The relevant questions include which hypotheses belong to the same inferential family, which analyses were prespecified, what error criterion the researchers intended to control, and whether the multiplicity strategy matches the claims.
If many statistical tests are central to the study and you cannot determine whether their inferential treatment is defensible, specialist review may be useful, particularly when the conclusions depend on a small number of favorable findings.
The consequence of relying on the study should influence your threshold
You do not need the same level of statistical assurance for every use of research evidence.
If you are reading a paper to understand an emerging idea, unresolved technical uncertainty may be tolerable. If the study will determine whether you recommend an intervention, change institutional practice, make a clinical decision, build a policy argument, include a high-weight estimate in a meta-analysis, or invest substantial resources, the consequences of analytical error are larger.
That does not mean every consequential decision requires a statistician. It means the value of reducing unresolved statistical uncertainty increases with the stakes.
Seek the right expertise, not merely someone who knows statistics
Statistics is not a single interchangeable specialty. A researcher who routinely performs conventional regression may not be the right person to evaluate a complex Bayesian hierarchical model. A machine-learning specialist may not be the ideal reviewer of a survey-weighted causal analysis. A psychometrician may be particularly valuable when the central problem concerns latent constructs or measurement models.
The ASA's ethical guidance recognizes that statistical practice occurs across diverse disciplines and emphasizes professional competence and appropriate collaboration. A useful consultation therefore matches the expert to the methodological problem rather than treating "statistician" as a universal troubleshooting category.
| Unresolved problem |
Expertise that may be particularly relevant |
| Complex causal effect estimation |
Statistician or epidemiologist specializing in causal inference |
| Prediction model development and validation |
Statistician or methodologist specializing in prediction modelling or machine learning |
| Latent constructs and measurement models |
Psychometrician, quantitative psychologist, or relevant statistician |
| Complex survey samples and weights |
Survey statistician |
| Bayesian hierarchical analysis |
Statistician experienced in Bayesian modelling |
| Longitudinal, clustered, or multilevel data |
Statistician experienced in longitudinal or multilevel modelling |
A good consultation question is specific
Sending someone a 30-page paper with the message "Can you check the statistics?" is possible, but it is not the most efficient use of expertise.
Instead, explain what you need to decide and where your uncertainty lies. For example: Does this interaction analysis support the authors' subgroup claim? Is the validation strategy likely to produce optimistic performance? Does adjustment for this post-treatment variable change the causal estimand? Is this confidence interval consistent with the model described?
A focused question gives the specialist a clearer target and helps you integrate the answer into the broader critical appraisal.
Statistical consultation does not transfer responsibility for interpretation
An expert may determine that the analysis is technically defensible. You still need to judge whether the outcome matters, whether the study population is relevant, whether the effect is substantively important, and whether the evidence applies to your question.
Statistical validity is one component of evidential credibility, not the whole of it.
The SAMPL guidelines themselves emphasize statistical reporting rather than replacing design-specific reporting guidance, while EQUATOR maintains separate guidelines for randomized trials, observational studies, systematic reviews, prediction research, diagnostic studies, and other designs. Statistical analysis and research design remain tightly connected.
Watch Out
Do not use consultation merely to obtain an authoritative-sounding endorsement. A useful statistical review should clarify assumptions, limitations, alternative interpretations, and unresolved uncertainty rather than reduce a complex appraisal to "the statistician said it was fine."
04 · A Practical Example
When Is Specialist Review Worth the Effort?
Hypothetical Example
A study may determine whether a university adopts an intervention
Suppose a university is considering an expensive student-support program. The main evidence comes from an observational study using inverse-probability weighting to estimate the program's effect. The crude analysis suggests little difference, while the weighted analysis reports a substantial benefit. The institution is considering using this study as a major justification for adoption.
Appraise what you can. You examine the population, outcome measurement, timing, attrition, intervention definition, reported effect magnitude, confidence interval, and whether the conclusion matches the reported weighted estimate.
Identify the unresolved issue. The main conclusion depends almost entirely on the weighting procedure. You understand its broad purpose but cannot confidently evaluate the propensity model, extreme weights, positivity, balance diagnostics, or sensitivity to unmeasured confounding.
Ask whether the uncertainty is central. If the weighting strategy is seriously flawed, the main estimate could change substantially. The unresolved statistical issue therefore sits directly beneath the study's central claim.
Consider the consequence. The evidence may influence a costly institution-wide decision. Reducing uncertainty has practical value beyond academic curiosity.
Seek matched expertise. A statistician or epidemiologist experienced in causal inference and propensity-based methods would be more useful than someone whose expertise lies primarily in unrelated statistical areas.
Ask focused questions. Request evaluation of the weighting specification, diagnostics, effective information after weighting, overlap, robustness, and whether the resulting estimate supports the causal interpretation claimed.
Integrate rather than outsource. Use the specialist assessment alongside your appraisal of design quality, outcome relevance, effect magnitude, feasibility, costs, and applicability to the institution.
In this situation, specialist review is valuable not because inverse-probability weighting sounds advanced, but because an unresolved technical issue is central to a consequential decision.