01 · The Question
What if the average effect does not represent everyone?
A meta-analysis reports a modest average benefit. Looking more closely, however, the intervention appears highly beneficial for one group, ineffective for another, and perhaps harmful for a third.
Should the overall effect still be your main conclusion?
Possibly, but not automatically. Average effects summarize a population. They do not prove that every individual or subgroup experiences the same relative or absolute effect. Genuine effect modification can matter greatly for practice and policy.
The difficulty is that subgroup analyses are also unusually good at producing persuasive patterns by chance. A responsible synthesis therefore has to take heterogeneity seriously without treating every subgroup difference as real.
03 · What You Need to Know
An average effect is a summary, not a promise of uniformity
What does an average effect actually tell you?
Suppose a meta-analysis estimates that an intervention reduces an outcome by an average of 15% relative to the comparator. That summary describes the effect represented by the included evidence under the model used.
It does not establish that every participant experiences a 15% reduction. Nor does it prove that the true relative effect is identical across age groups, baseline risk levels, settings, intervention variants, or other characteristics.
Variation in effects can arise because an intervention genuinely works differently under different conditions. This is often described as effect modification or heterogeneity of treatment effects.
Overall effect
A summary of the intervention contrast across the population or studies represented in the analysis.
Subgroup effect or effect modification
A difference in the intervention effect according to a characteristic such as population, setting, intervention feature, or another prespecified factor.
Different absolute effects do not necessarily mean different relative effects
This distinction is crucial.
Imagine that an intervention reduces relative risk by approximately the same proportion in both high-risk and low-risk groups. Because the high-risk group begins with more events, the same relative reduction can produce a much larger absolute benefit in that group.
That is not necessarily evidence that the intervention has a different relative effect across subgroups. It may simply reflect different baseline risks.
For decisions, those differences in absolute benefit can still be extremely important. Methodological work on subgroup interpretation has emphasized that prognostic differences can create meaningful variation in absolute treatment effects even when relative effects remain similar.
| Pattern |
Possible interpretation |
What to examine |
| Similar relative effect, different baseline risks |
Absolute benefits may differ without relative effect modification |
Baseline risk and absolute effects |
| Different relative effects across subgroups |
Possible effect modification |
Interaction evidence and subgroup credibility |
| Different effects across separate studies |
Could reflect effect modification or other study differences |
Within-study evidence, confounding, and study characteristics |
| One subgroup significant, another non-significant |
Does not establish a subgroup difference |
Direct statistical comparison of effects |
“Significant here, not significant there” is not a test of subgroup difference
This is perhaps the most common subgroup error.
Suppose the intervention effect is statistically significant among younger participants but not among older participants. It does not follow that age modifies the intervention effect.
The correct question is whether the estimated effects differ from each other beyond what would reasonably be expected from sampling variation. That requires an appropriate interaction test or equivalent direct assessment of the difference between subgroup effects.
Methodological guidance consistently warns against comparing the statistical significance of separate subgroup estimates. Reviews should report how subgroup effects were compared, including whether an interaction test was used.
Watch Out
“Significant in subgroup A but not in subgroup B” does not mean “significantly different between A and B.” The latter requires a direct comparison of the subgroup effects.
Subgroup analyses multiply opportunities for chance findings
If researchers divide participants by age, sex, baseline severity, prior treatment, location, socioeconomic status, adherence, and several other characteristics, some apparently striking differences may arise simply through repeated testing.
The more subgroup hypotheses examined after seeing the data, the easier it becomes to find a compelling story.
Credibility is therefore stronger when the subgroup hypothesis was specified in advance, when relatively few hypotheses were tested, and when the expected direction of effect was prespecified. Methodological assessments have found that subgroup claims in published trials often fail several established credibility criteria.
Within-study subgroup comparisons are generally more credible than between-study comparisons
Suppose you suspect that age modifies an intervention effect.
The strongest evidence would ideally compare younger and older participants within the same randomized studies. Because both subgroups arise within the same studies, many study-level characteristics are held constant.
A weaker approach might compare studies whose participants have an average age of 35 with entirely different studies whose participants have an average age of 70. Those studies may differ in setting, intervention implementation, eligibility criteria, follow-up, comparator, or numerous other characteristics.
An apparent relationship between average age and treatment effect could therefore reflect another difference between the studies. Contemporary GRADE guidance likewise regards within-study evidence of effect modification as substantially more compelling than comparisons based only on differences between studies.
Subgroups defined after treatment begins can be especially problematic
Subgroups based on baseline characteristics such as age or pre-intervention severity can be considered without being consequences of the assigned treatment.
Post-intervention characteristics are different. Suppose researchers divide participants according to whether they adhered well to treatment, completed the intervention, or remained in care for a particular duration. Treatment itself may influence those characteristics.
Comparisons based on such post-randomization variables can destroy the original comparability produced by randomization and generate misleading subgroup conclusions. Established credibility criteria therefore give much greater confidence to subgroup variables defined at baseline.
A plausible explanation helps, but does not prove the subgroup effect
A subgroup hypothesis is more credible when supported by prior evidence or a convincing rationale. For example, researchers may have strong reasons established before the study to expect an intervention to behave differently at different levels of baseline severity.
But plausibility cannot rescue weak empirical evidence. A compelling explanation constructed after an unexpected subgroup result can be particularly seductive because humans are quite good at explaining patterns once we already know they occurred.
This is another version of the broader problem of treating mechanistic plausibility as proof of an effect.
Consistency across studies strengthens subgroup claims
If several independent studies repeatedly show a similar interaction, the subgroup hypothesis becomes more credible than if it appeared once in a small analysis.
Consistency should concern the subgroup interaction itself, not merely whether the intervention works in the same subgroup repeatedly.
For example, showing benefit among younger participants in several studies does not establish that younger participants benefit more than older participants unless the evidence actually supports a difference in treatment effects between those groups.
Subgroup effects exist on a credibility continuum
No single criterion proves that a subgroup effect is real. Methodological frameworks therefore recommend considering several features together: whether the hypothesis was prespecified, whether the subgroup characteristic was measured at baseline, whether only a limited number of hypotheses were tested, whether an appropriate interaction analysis supports the difference, whether the effect is consistent across studies or related outcomes, and whether external evidence supports the hypothesis.
A subgroup finding satisfying many of these considerations deserves more confidence than an isolated post hoc finding among dozens of exploratory analyses.
The goal is not to ban exploratory subgroup analysis. Exploratory findings can generate valuable hypotheses. The problem begins when hypothesis-generating evidence is presented as though it were already confirmatory.
04 · A Practical Example
How can an average effect hide a potentially important subgroup difference?
Hypothetical Example
A digital tutoring intervention appears modestly effective overall
Imagine a hypothetical meta-analysis of randomized trials evaluating an adaptive tutoring system. The overall result suggests a modest improvement in examination performance. Researchers hypothesized before the review that the intervention might work differently for students according to baseline academic performance.
Start with the overall effect The reviewer reports the pooled average effect rather than discarding it merely because subgroup variation is suspected.
Test the subgroup hypothesis directly Studies providing results for lower- and higher-performing students are used to examine whether the intervention effect differs between those groups. The reviewer focuses on the interaction rather than asking whether each subgroup separately reaches statistical significance.
Assess credibility The hypothesis was prespecified, baseline performance was measured before intervention, only a small number of subgroup hypotheses were planned, and several trials provide within-study comparisons.
Examine consistency The reviewer checks whether the direction of the interaction is similar across studies rather than allowing one particularly dramatic trial to determine the conclusion.
Calibrate the conclusion If the interaction is supported but still imprecise, the review may conclude that the intervention appears to produce larger effects among students with lower baseline performance while acknowledging uncertainty about the magnitude of that difference.
The subgroup analysis has not replaced the overall effect. It has refined the question by asking whether the average conceals systematic variation that is sufficiently credible to matter.
06 · What This Means for You
Ask whether the subgroup difference itself is credible
When a pooled effect appears to hide important variation, do not jump directly from subgroup-specific estimates to subgroup-specific conclusions. Evaluate the evidence for effect modification explicitly.
A simple decision framework
If subgroup effects were hypothesized before the results were known
Give the analysis more consideration, particularly when the hypothesis was one of a limited number and its expected direction was specified.
If the apparent difference comes from significance in one subgroup but not another
Do not infer effect modification; examine a direct test of the difference between subgroup effects.
If subgroup evidence comes from comparisons within the same studies
Generally regard it as more informative about effect modification than a relationship based only on differences between separate studies.
If a subgroup pattern emerged from many exploratory analyses
Treat it primarily as hypothesis-generating unless independent evidence provides convincing confirmation.
If relative effects are similar but baseline risks differ
Present subgroup-specific absolute effects where useful rather than incorrectly claiming different relative treatment effects.
When reporting the synthesis, describe how subgroup variables were defined, whether hypotheses were prespecified, whether comparisons were within or between studies, how many subgroup hypotheses were investigated, and how the subgroup effects were statistically compared. Cochrane reporting guidance specifically calls for these details when subgroup analyses or meta-regression are used.
Above all, match the strength of your language to the credibility of the subgroup evidence. “The intervention is effective only for subgroup A” is a much stronger claim than “the available evidence suggests that effects may be larger in subgroup A.” The data have to earn the stronger sentence.
07 · A Quick Checklist
Before concluding that average effects hide a subgroup difference, check:
For each proposed subgroup effect, verify:
Determine whether the subgroup hypothesis was specified before the results were known.
Prefer subgroup characteristics measured at baseline rather than variables potentially affected by the intervention.
Check how many subgroup hypotheses were tested and whether the analysis was confirmatory or exploratory.
Use an interaction test or appropriate equivalent rather than comparing separate P values across subgroups.
Distinguish within-study subgroup comparisons from less reliable comparisons based only on differences between studies.
Examine whether the subgroup interaction is consistent across studies and supported by relevant prior evidence.
Distinguish genuine relative effect modification from differences in absolute benefit caused by different baseline risks.
Present exploratory subgroup findings with appropriate uncertainty rather than converting them into definitive treatment rules.