03 · What You Need to Know
How to identify the outcome behind a reported result
Start with the variable, not the conceptual label
An outcome is often described at two levels. The first is the concept researchers care about, such as depression, academic achievement, mobility, treatment response, or quality of life. The second is the variable through which that concept becomes observable in the study.
For example, “depression” might be represented by a score on a specified questionnaire, a diagnostic interview, remission status, hospitalization, medication use, or another indicator. Those measures answer related but different questions.
Outcome construct
The broader concept or state the researchers want to understand.
Outcome measure
The specific variable, event, score, classification, or observation used to represent that outcome empirically.
Critical appraisal requires you to keep the distinction visible. Otherwise, a result about one measurement can quietly become a claim about an entire construct.
A complete outcome requires more than naming the measurement
Contemporary trial guidance illustrates how much information can be hidden inside a seemingly simple outcome. SPIRIT 2025 recommends specifying, for each primary and secondary outcome, the measurement variable, participant-level analysis metric, method of aggregation, and time point.
Consider the phrase “depression score.” It remains incomplete. Is the outcome the final score at follow-up, change from baseline, percentage change, proportion crossing a response threshold, or time until remission? The same underlying measurements can be transformed into different analytical outcomes.
| Outcome component |
Question to ask |
Example |
| Measurement variable |
What was directly recorded? |
Score on a specified depression scale |
| Analysis metric |
What participant-level value entered the analysis? |
Change from baseline |
| Aggregation |
How was the outcome summarized across participants? |
Mean change in each group |
| Time point |
When was the outcome of interest evaluated? |
12 weeks after randomization |
SPIRIT's explanation notes that the same measurement variable can support metrics such as change from baseline, final value, or time to an event. Patient-reported-outcome guidance similarly shows that several analytical endpoints can be derived from the same underlying outcome data.
Find out exactly how the outcome was measured
If the outcome is a score, identify the instrument. If it is a diagnosis, identify the diagnostic criteria and assessment procedure. If it is an event, determine what counted as the event and how it was ascertained. If it comes from administrative data, identify the relevant data source and coding definition.
Measurement method matters because two instruments bearing similar labels may capture different domains, use different scales, or have different reliability and validity for the population being studied.
Also identify who provided or assessed the outcome when relevant. A self-reported symptom, clinician-rated symptom, parent report, teacher rating, device measurement, and administrative record may provide different perspectives on ostensibly similar phenomena.
Outcome labels can hide composite definitions
Some outcomes combine several events into one composite. A cardiovascular composite, for example, might count the first occurrence of cardiovascular death, myocardial infarction, or stroke. Another study could use the same broad label but include hospitalization as an additional component.
When interpreting a composite outcome, identify every component and determine what event actually counted toward the result. A statistically detectable composite effect can sometimes be driven more strongly by particular components than others.
Binary outcomes require you to inspect the threshold
Continuous measurements are often converted into categories such as responder versus non-responder, recovered versus not recovered, or high versus low risk.
Ask how the threshold was chosen. For example, a “successful outcome” might mean a score below a specified cutoff, improvement by a certain number of points, or a percentage reduction from baseline.
Changing the threshold can change who counts as having experienced the outcome, even when the underlying measurements remain identical.
Change from baseline and final value are different outcome metrics
Suppose students take the same achievement test before and after an intervention. Researchers could analyze final test scores, absolute change from baseline, percentage change, or another metric.
These analyses use related data but are not simply different names for the same outcome. You should identify which metric corresponds to the result being interpreted.
SPIRIT 2025 explicitly includes final value, change from baseline, and time to event as examples of different participant-level analysis metrics that should be specified.
Time-to-event outcomes require both an event and a clock
For outcomes such as survival, relapse, hospitalization, or treatment discontinuation, knowing what event occurred is not enough. The analysis also depends on when follow-up begins, what constitutes the event, how censoring is handled, and the period over which participants are observed.
“Mortality” might mean death during hospitalization, 30-day all-cause mortality, one-year disease-specific mortality, or time to death over the complete follow-up period. These are different outcomes.
Primary and secondary outcomes should remain distinct
Studies frequently measure several outcomes. In trials, the outcome of main interest is generally designated as the primary outcome, while other outcomes may be secondary or exploratory. SPIRIT 2025 notes that the primary outcome usually corresponds closely to the study objective and commonly informs the sample size calculation.
A paper may nevertheless devote substantial attention to an interesting secondary result. Do not let prominence in the discussion retroactively turn a secondary outcome into the primary outcome.
Determining what was prespecified and what appears to have been decided later can be important when the hierarchy of outcomes is unclear.
Surrogate outcomes are not automatically equivalent to outcomes that matter directly
Researchers sometimes measure biomarkers, physiological indicators, intermediate behaviors, test scores, or other outcomes because they are easier or faster to observe than more distal outcomes.
A surrogate or intermediate outcome may be scientifically useful, but improvement in it does not automatically establish improvement in every downstream outcome it is intended to represent. The strength of that inference depends on the context and evidence linking the surrogate to the outcome of interest.
For example, an educational intervention that improves performance on an immediate researcher-developed task has demonstrated an effect on that task. Whether it improves durable learning, transfer to unfamiliar problems, later course performance, or professional competence remains a separate empirical question unless those outcomes were also measured.
Patient-reported and self-reported outcomes are real outcomes, but their source matters
An outcome does not become invalid merely because participants report it themselves. Symptoms, experiences, functioning, satisfaction, and quality of life may be appropriately assessed through patient- or participant-reported measures.
The relevant questions concern what the instrument measures, how it is scored, whether it is appropriate for the population and purpose, how missing responses are handled, and what interpretation the score supports. Guidance for patient-reported outcomes emphasizes specifying the relevant domain, instrument, analysis metric, and time point.
The measured outcome may be narrower than the conclusion
This is one of the most useful checks in critical reading.
If researchers measured examination performance, a conclusion about “learning” may require qualification. If they measured intentions to vaccinate, they did not directly measure vaccination behavior. If they measured self-reported productivity, they did not necessarily measure objectively recorded output. If they measured a biomarker, they did not automatically measure clinical benefit.
Watch Out
Whenever the conclusion uses a broader noun than the methods, compare the two carefully. Ask whether the study directly measured the broader outcome or whether the authors are making an inferential step from an operational indicator to a larger construct.
One paper can contain many versions of an outcome
A study might assess the same construct with several instruments, evaluate it repeatedly, report continuous and dichotomized versions, or analyze both change scores and final values.
Do not ask merely “What outcomes did they collect?” Ask which outcome definition generated the specific estimate you are reading. Then inspect the primary analysis to see how that outcome entered the model.