03 · What You Need to Know
You Can Evaluate More Than You May Think Without Pretending to Be a Specialist
Start by separating unfamiliarity from evidence of error
An unfamiliar method is not automatically inappropriate. Modern research questions often require methods beyond introductory t tests, correlations, ordinary regression, and conventional analysis of variance.
Repeated observations may require models that account for dependence. Hierarchical data may require multilevel methods. Missing data may require methods more defensible than deleting incomplete cases. Time-to-event outcomes need analyses that account for censoring. Prediction models may require penalization and careful validation.
Complexity can therefore be justified by the problem.
The reverse is also true. A sophisticated method is not automatically correct merely because it is difficult to understand. Technical vocabulary should not function as methodological diplomatic immunity.
I do not understand this method
A statement about the limits of your present statistical expertise.
This method is inappropriate
A methodological conclusion that requires evidence about the method, assumptions, implementation, or fit to the research question.
Ask what problem the method is supposed to solve
You do not need to derive a model mathematically before asking why the researchers used it.
Identify the structure of the problem. Are observations clustered within schools, hospitals, countries, families, or participants? Are measurements repeated over time? Is the outcome binary, ordinal, count-based, continuous, or time-to-event? Are there missing data? Is the study estimating a causal effect, describing an association, or predicting future outcomes?
Then ask whether the paper explains why the chosen method is appropriate for that structure.
If authors use a multilevel model because students are nested within classes and schools, you can understand the methodological purpose even if you could not fit the model yourself. If they use an elaborate model without explaining what feature of the data or research question requires it, appraisal becomes harder.
Separate the statistical model from the scientific design
Advanced analysis cannot repair a fundamentally weak design.
You can still ask whether exposure preceded outcome, whether comparison groups are credible, whether selection could bias the sample, whether variables were measured appropriately, whether important confounding remains, whether missing data are substantial, and whether outcomes were selectively reported.
A sophisticated causal model applied to poorly measured observational data still inherits limitations from those data. An elegant hierarchical model cannot make an inappropriate outcome measure valid. A machine-learning algorithm cannot turn leakage between training and test data into genuine predictive performance.
Do not allow mathematical complexity to push basic design appraisal off the page.
Check whether the method is described well enough to evaluate or reproduce
Reporting quality is something you can assess even when the method itself is specialized.
CONSORT 2025 states that statistical procedures and software should be specified and that methods should be described in sufficient detail for a knowledgeable reader with access to the original data to verify the reported results. SAMPL likewise provides guidance for reporting statistical methods and analyses.
Look for the model type, outcome, predictors or covariates, interactions, transformations, estimation approach, missing-data procedures, statistical software, uncertainty measures, and relevant model diagnostics or sensitivity analyses.
If critical analytical decisions are absent, the problem is not merely that you lack expertise. The paper may not provide enough information even for an expert to evaluate the analysis properly.
Look for assumptions you can identify
Every statistical method rests on assumptions, although they differ considerably across methods. You may be able to identify some of them without mastering the complete theory.
| Method or problem |
Questions you may be able to ask |
| Regression model |
Were important relationships modeled plausibly? Were influential observations or functional forms considered? |
| Multilevel model |
Does the data structure actually contain meaningful clustering? Are relevant levels represented? |
| Survival analysis |
How was censoring handled? If proportional hazards are assumed, was that assumption considered? |
| Missing-data method |
How much data were missing? What assumptions about missingness are required? Were sensitivity analyses performed? |
| Prediction model |
Was performance evaluated on data not simply reused to build the model? Were calibration and discrimination assessed? |
| Bayesian analysis |
What priors were used? Were they justified? Are sensitivity analyses to plausible prior choices available where relevant? |
You may not be able to judge every diagnostic, but these questions help determine exactly where specialist evaluation is required.
Examine whether the results are presented transparently
Even when the underlying model is complicated, the substantive results should usually be interpretable.
What quantity was estimated? In what direction? How large was it? How uncertain is it? What population or comparison does it refer to?
If a paper reports only model coefficients whose substantive meaning is opaque, look for marginal effects, predicted probabilities, contrasts, risk differences, or other interpretable summaries where appropriate. Complexity in estimation does not require obscurity in communication.
The principles behind interpreting effect estimates and uncertainty rather than relying only on P-values remain relevant even when the model producing those quantities is advanced.
Check whether conclusions depend on one fragile analytical specification
Sensitivity analyses can be particularly useful when you cannot independently verify every mathematical detail. Ask whether reasonable alternative assumptions, models, variable definitions, adjustment strategies, missing-data approaches, or priors produce substantively similar conclusions.
Consistency across defensible alternatives does not prove the preferred analysis correct. It can, however, show that the central conclusion is not entirely dependent on one narrow analytical choice.
If small modelling changes reverse the conclusion, that fragility deserves attention.
Check whether adjustment choices make conceptual sense
You may not be able to derive the estimator, but you may still be able to evaluate the variables entering it.
For example, if a complex causal analysis adjusts for a variable measured after the exposure, you can ask whether that variable could be a mediator or otherwise affected by treatment. If dozens of covariates were selected according to observed P-values, you can question the selection rationale.
The principles for judging whether statistical adjustment is appropriate remain relevant regardless of whether the final estimator is simple or mathematically sophisticated.
Do not substitute peer review for statistical understanding
Publication in a peer-reviewed journal provides useful context, but peer review does not guarantee that every statistical choice has been examined by a specialist with the relevant expertise.
The paper may have received statistical review, or it may not. Even specialist review cannot guarantee correctness.
Treat peer review as one part of the publication process, not as a certificate allowing you to skip methodological appraisal.
Do not substitute software for justification either
The fact that a method is implemented in R, Stata, SAS, SPSS, Python, or another established package does not show that it was appropriate for the data.
Software can execute a model exactly as requested while the requested model is scientifically inappropriate. Correct syntax is not the same as correct inference. Somewhere, a perfectly converged model is quietly answering the wrong question.
Know when the unresolved uncertainty actually matters
Not every technical detail requires specialist consultation.
If the advanced analysis concerns a peripheral exploratory outcome and does not affect your use of the study, documenting your uncertainty may be sufficient. If the entire paper's main claim depends on a complex model whose assumptions you cannot assess, the situation is different.
Ask what would happen if the analysis were wrong. Would the central conclusion collapse? Would a clinical, educational, policy, financial, or research decision change? Would the study receive substantial weight in a systematic review?
The more consequential the dependence on the method, the stronger the case for seeking statistical expertise before relying heavily on the study.
Statistical expertise is not one generic skill
A statistician experienced in randomized clinical trials may not specialize in Bayesian computation, causal inference, psychometrics, complex survey estimation, spatial statistics, machine learning, or longitudinal latent-variable modeling.
When seeking help, match expertise to the method and research design rather than assuming that any quantitatively trained person can evaluate every advanced analysis.
The American Statistical Association's ethical guidance recognizes that statistical practitioners have different expertise and experience and emphasizes appropriate professional consultation. It also stresses transparency about assumptions, limitations, methods, and possible sources of error.
Watch Out
Do not write "the statistical analysis appears appropriate" merely because you cannot identify an error. If you cannot evaluate a consequential part of the analysis, say what you could assess and what remains uncertain.