01 · The Question
Should Critical Appraisal Happen While You Read or After You Finish the Literature?
You are reading a study and immediately notice something consequential: perhaps the sampling strategy limits generalizability, the measurement approach is unusually strong, or the design cannot support the causal language used in the discussion. Should you record that judgment now, or wait until you have read enough literature to evaluate the study properly?
Both instincts have merit. Waiting can give you valuable comparative context. But if you postpone all evaluation, you may forget the reasoning that was obvious while the methods, results, and authors' claims were directly in front of you.
The useful distinction is between evaluating features of the individual study and judging what those features mean within the larger body of evidence.
02 · The Short Answer
Record Them While Reading, Then Reassess Them Later
In Brief
Record study-specific strengths and weaknesses while reading, especially when they affect how you should interpret the findings, but treat those judgments as provisional and revisit them during synthesis.
You can often evaluate design choices, measurement, sampling, analysis, and the match between evidence and conclusions from the paper itself. Their broader significance, however, may become clearer only after you compare the study with others addressing the same question.
03 · What You Need to Know
Critical Appraisal Works at More Than One Stage
Critical reading does not mean searching every paper for flaws. It means asking how trustworthy, informative, and applicable its evidence is for the question you are examining. Critical appraisal commonly involves examining study design, methods, results, potential sources of bias, applicability, and the relationship between the evidence and the conclusions drawn from it.
That work can begin while you read. It does not have to end there.
Some judgments are easiest to make while the paper is open
Suppose you notice that a study uses a validated measure, clearly describes participant recruitment, reports substantial attrition, relies entirely on self-reported outcomes, or interprets a cross-sectional association causally. These observations depend primarily on information contained in that study.
Recording them immediately has a practical advantage: the evidence supporting your judgment is still visible. You can also record the relevant page, table, passage, or methodological detail rather than trusting yourself to reconstruct the criticism months later.
University guidance on literature reviews commonly recommends recording methodology, findings, implications, and limitations during structured reading because these notes later provide material for analysis and synthesis.
Not every weakness is a fatal flaw
A limitation is meaningful in relation to what a study is trying to establish. A small purposive sample, for example, may be problematic for a claim about population prevalence but entirely defensible for a qualitative inquiry seeking detailed accounts of a particular experience.
Similarly, the absence of random assignment is not automatically a defect. Many research questions cannot ethically or practically be investigated through randomized experiments. The appropriate question is not simply, “What didn't the researchers do?” It is, “Given this research question and design, what can this evidence reasonably support?”
Study feature
An observable characteristic of the research, such as its sampling procedure, design, measurement approach, comparison group, analytic method, or follow-up period.
Evaluative judgment
Your reasoned assessment of how that feature strengthens, limits, qualifies, or otherwise affects the interpretation of the evidence.
Separating observation from judgment can prevent vague criticism. “Sample size: 42” is an extracted fact. “The sample may be insufficient for the precision required by this analysis” is an evaluation that requires justification.
Some weaknesses become visible only through comparison
You may initially consider a study's measurement approach perfectly reasonable. After reading twenty additional studies, however, you discover that most use direct performance measures while this one relies on a single self-report item. That difference may become important when explaining divergent findings.
The reverse can happen too. Something that initially looks like a weakness may turn out to be normal or necessary within that research tradition.
This is why synthesis involves more than evaluating papers individually. Researchers also need to examine patterns, contradictions, methodological differences, and gaps across the literature.
Record observations before they disappear into a general impression
One danger of postponing evaluation is that specific methodological observations gradually collapse into labels such as “strong paper” or “weak study.” Those labels are rarely useful.
Instead, record the reason:
Outcome measured with a validated instrument rather than a single researcher-created item.
Longitudinal design establishes temporal ordering more clearly than the cross-sectional studies in this theme.
Large amount of attrition between baseline and follow-up may affect interpretation.
Recruitment from one specialized institution may limit applicability to other settings.
Authors acknowledge alternative explanations but the design cannot distinguish among them.
These notes preserve the analytical basis of your judgment. They are also easier to revise when your understanding changes.
Distinguish author-reported limitations from your own appraisal
A paper's limitations section is useful, but it is not a substitute for critical appraisal. Authors may identify some limitations and omit others. Conversely, researchers sometimes describe features as limitations that may be relatively inconsequential for the question you are asking.
Your notes should therefore make the source of each observation clear. This connects directly to the need to separate what the authors said from your own interpretation .
For example:
Authors' limitation: “The authors note that the study was conducted at one institution.”
My appraisal: “Single-institution recruitment may matter here because institutional AI policies differ substantially across settings, making transfer of the finding uncertain.”
The second statement is your interpretation. It should not later migrate into your notes as though the authors themselves made that argument.
Strengths deserve the same specificity as weaknesses
Critical appraisal can easily become flaw collecting. That is not its purpose. A study may contribute unusually strong measurement, an informative comparison group, careful handling of confounding, transparent reporting, an appropriate design, or evidence from an underrepresented context.
A balanced appraisal asks what the evidence can support as well as what it cannot. Literature reviews are expected to critically examine strengths, weaknesses, and omissions rather than merely summarize studies.
Your final evaluation should emerge from comparison
During synthesis, return to your initial appraisals and ask whether they still matter. Perhaps five studies reach similar conclusions despite using substantially different methods. Perhaps apparently conflicting findings separate neatly according to measurement approach. Perhaps nearly every study shares the same limitation.
At this point, the analytical unit begins shifting from individual papers to patterns across studies. That is why recording why individual papers matter can help: methodological strengths and weaknesses become especially useful when they explain the contribution or evidential weight of a source.
07 · A Quick Checklist
What to Check When Appraising a Study
While reading and again during synthesis, check:
Have I recorded concrete methodological features rather than vague labels such as “good” or “weak”?
Have I considered whether the design is appropriate for the question the study attempts to answer?
Have I examined sampling, measurement, analysis, missing data, and other issues that materially affect interpretation when relevant?
Have I distinguished author-reported limitations from my own critical assessment?
Have I recorded strengths as carefully as limitations?
Have I avoided criticizing a study merely for not using a design inappropriate to its research question?
Have I revisited my initial appraisal after comparing the study with related evidence?
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation