01 · The Question
Can evidence from children, younger adults, and older adults really be combined?
Your review question concerns a broad population, but the available studies do not divide neatly by age. One study includes adolescents. Another includes adults aged 18 to 65. A third focuses on people over 60. Several report only the mean age of participants.
Should these studies contribute to one synthesis, or does combining them risk hiding age-related differences?
Age can matter substantially, but an age difference does not automatically make evidence incomparable. The important question is whether age could plausibly change the outcome, effect, mechanism, measurement, or applicability relevant to your review.
03 · What You Need to Know
Age can matter in several different ways
Age is often treated as a routine demographic variable, yet its methodological meaning depends on the research question. It may represent biological development, accumulated exposure, disease risk, social roles, educational stage, cognitive development, treatment tolerance, technology use, or numerous other characteristics.
That makes “Does age matter?” too broad a question. You need to ask how age could matter for the specific phenomenon being synthesized.
Age difference does not automatically create indirect evidence
GRADE guidance treats population differences, including age differences, as a potential source of indirectness when researchers apply evidence to a target population. Yet differences between the studied and target populations do not automatically justify lower confidence. The concern becomes important when there is a credible reason to expect those differences to change the effect relevant to the target question.
For example, evidence from adults may sometimes inform decisions about older adults if the underlying mechanism and relative effect are expected to remain similar. In another context, evidence from adults may provide poor guidance for young children because developmental physiology, dosing, measurement, behavior, or implementation differs fundamentally.
The substantive question determines which situation you face.
Distinguish baseline risk from effect modification
An age group can have a different baseline probability of an outcome without necessarily experiencing a different relative intervention effect.
Suppose an adverse event is much more common among older adults. Even if an intervention produces the same relative risk reduction across ages, the absolute benefit may be larger among older adults because their starting risk is higher.
Prognostic effect
Age is associated with the likelihood of the outcome regardless of the intervention or exposure.
Effect modification
The effect of the intervention or exposure itself differs according to age.
Confusing these concepts can lead researchers to conduct age subgroup analyses simply because age predicts the outcome. A prognostic factor is not necessarily an effect modifier.
Do not create age categories mechanically
Age is fundamentally continuous, even though studies frequently convert it into categories such as children, adolescents, younger adults, middle-aged adults, and older adults. Those categories may be substantively useful, but their boundaries should not be treated as natural laws.
A threshold of 60 or 65 years may make sense for one research question and little sense for another. Developmental research may require much narrower categories, while another phenomenon may change gradually across adulthood.
When possible, base age categories on biological, clinical, developmental, social, or policy relevance rather than selecting cut points merely because they make subgroup analysis convenient.
Broad age ranges can hide within-study variation
A study reporting a mean participant age of 45 years might include people aged 18 to 80. Another study with the same mean could contain almost exclusively participants aged 40 to 50. Their reported means therefore do not establish equivalent age composition.
This becomes particularly important in meta-regression. Cochrane guidance highlights age as an example of aggregation bias: a true relationship between age and treatment effect may exist within individual studies but remain invisible when researchers compare only the mean ages of entire studies.
The reverse problem can also occur. A relationship between study-level mean age and effect estimates does not prove that age modifies the effect among individuals.
Watch Out
Do not interpret a study-level meta-regression of mean age as if it were an individual-participant analysis. Study averages discard within-study information and can produce ecological or aggregation bias.
Ask whether the intervention operates differently by age
Age deserves greater attention when there is a plausible reason that the intervention or exposure would operate differently across the age range.
In clinical research, pharmacokinetics, comorbidities, organ function, developmental physiology, treatment tolerance, or competing risks may matter. In education, developmental stage, prior knowledge, autonomy, and curriculum can change the meaning of an intervention. In technology research, device familiarity or patterns of use may differ across generations.
None of these differences should simply be presumed. They provide reasons to investigate whether age-related effect modification is plausible.
Check whether the outcome has the same meaning across ages
Measurement comparability can become as important as intervention comparability. The same instrument may not have equivalent validity across children and adults. A behavioral outcome may have different developmental meanings. Functional independence, academic performance, social participation, or quality of life may also be operationalized differently across life stages.
If outcomes are not measuring sufficiently comparable constructs, pooling them because they share a label can produce a deceptively precise answer to an unclear question.
Compare effects directly rather than comparing significance
A familiar subgroup mistake occurs when researchers observe a statistically significant effect among younger participants and a nonsignificant effect among older participants and conclude that the treatment works only in younger people.
That conclusion does not follow. Statistical significance depends partly on sample size and precision. The effect estimates themselves might be very similar.
To evaluate age-related effect modification, the relevant analysis examines evidence for a difference between effects. Even then, subgroup findings should be interpreted alongside their precision, prespecification, biological or substantive plausibility, consistency across studies, and vulnerability to confounding or multiple testing.
Sometimes individual participant data provide a better answer
When age-related modification is central to the research question, aggregate study-level data may be inadequate. Individual participant data meta-analysis can allow age to be modeled at the participant level rather than relying on study means or broad published categories.
This can be particularly useful when age is continuous or when different studies use incompatible age cutoffs. It does not remove every source of bias, but it can avoid some limitations inherent in study-level age comparisons.
Age also affects applicability
Even when a synthesis is statistically coherent, you still need to ask whether it applies to the population of interest. A body of evidence dominated by middle-aged adults may offer less direct evidence for children or very old adults when age-related differences could plausibly alter the finding.
This is part of the broader question of when evidence from another population is directly relevant to yours. The answer depends on the characteristics that matter to the phenomenon, not simply whether the populations have different demographic labels.
06 · What This Means for You
Let the research question determine how much age should structure the synthesis
Before dividing evidence into age groups, specify why age could change the finding. That reasoning should guide extraction, subgroup definitions, synthesis, and interpretation.
A simple decision framework
If age differs across studies but no plausible age-related effect modification is expected
A combined synthesis may remain appropriate, provided the studies are otherwise sufficiently comparable.
If developmental, biological, behavioral, or contextual mechanisms make age potentially important
Preserve age-specific information and prespecify appropriate analyses where possible.
If studies report usable age-specific effect estimates
Compare the estimates directly rather than relying on separate within-group significance tests.
If only study-level mean ages are available
Treat meta-regression cautiously and avoid translating study-level associations directly into individual-level conclusions.
If the evidence scarcely represents your target age group
Consider whether the evidence is indirect rather than silently assuming that effects generalize.
Sometimes the most useful conclusion is that the available evidence does not permit a reliable age-specific inference. That is not a failed synthesis. It identifies precisely where the evidence stops supporting confident interpretation.
If the target age group differs substantially from those studied and there is a credible reason the finding could change, it may be necessary to state explicitly that the evidence does not generalize directly.
07 · A Quick Checklist
Before combining evidence across age groups, check:
Before synthesizing across ages, check:
Does age have a plausible biological, developmental, behavioral, or contextual relationship with the finding?
Are you distinguishing age-related baseline risk from genuine effect modification?
Do the intervention, exposure, comparator, and outcome have comparable meanings across the age range?
Are age categories substantively justified rather than selected merely for analytical convenience?
Have you examined age ranges and distributions rather than relying only on study means?
Are subgroup conclusions based on direct comparisons between effects rather than differences in statistical significance?
Could study-level age analyses suffer from confounding, ecological bias, or aggregation bias?
Is your target age group adequately represented in the evidence?