03 · What You Need to Know
Diversity can make consistency more informative
Imagine five studies conducted in nearly identical hospitals, using similar participants, personnel, implementation procedures, and comparators. Now imagine five equally rigorous studies conducted across different health systems, resource environments, population groups, and implementation contexts.
If both sets produce similar effects, the second pattern can answer an additional question: does the finding appear dependent on the contextual characteristics that vary across those studies?
This does not mean heterogeneous settings are inherently better. Rather, variation creates an opportunity to observe whether the effect persists when potentially relevant conditions change.
First ask whether the settings are different in ways that matter
Geographical diversity alone provides limited information. Ten studies from ten countries might still involve highly similar tertiary hospitals, affluent participants, specialist personnel, and intervention protocols.
Conversely, two studies within one country may differ substantially in infrastructure, institutional organization, population characteristics, baseline risk, or implementation.
Therefore, “multicountry” and “contextually diverse” are not synonyms.
Geographical diversity
Studies occur in different locations or countries.
Effect-relevant contextual diversity
Studies vary on characteristics that could plausibly modify the effect or mechanism.
The second form of diversity is what makes consistency especially informative.
Consistency matters most when there was a credible reason to expect variation
Suppose researchers suspect that an intervention requires extensive institutional support. If comparable effects emerge in settings with substantially different levels of support, that observation challenges the hypothesis that the effect depends strongly on that characteristic.
Similarly, if an educational intervention performs comparably across different class sizes, delivery systems, and student populations, confidence may increase that its effect is not restricted to one particular implementation environment, assuming those differences were measured adequately.
This reasoning resembles a robustness test. The finding has encountered variation that could have disrupted it, yet the observed effects remain reasonably compatible.
Consistency does not require identical estimates
Real studies rarely produce precisely the same numerical effect. Sampling variation alone ensures some differences, and genuine contextual variation may remain even when effects broadly point in the same direction.
Cochrane distinguishes statistical heterogeneity from the broader clinical and methodological diversity among studies. An evidence base can show some variation while still supporting a broadly consistent substantive conclusion.
Consequently, the question is not whether every effect estimate is numerically identical. Ask whether the pattern remains sufficiently coherent given uncertainty, study differences, and the decision being made.
Do not reduce consistency to an I² value
A low heterogeneity statistic can be reassuring in some circumstances, but it does not establish cross-context robustness by itself.
First, statistical heterogeneity describes variation in observed effects, not how diverse the underlying settings are. Ten nearly identical studies can have low heterogeneity without testing contextual robustness at all.
Second, estimates of heterogeneity can be uncertain, especially when few studies are available. Third, effects can differ in practically important ways even when a formal heterogeneity test has limited power.
Contextual consistency therefore requires reading the studies, not merely reading the heterogeneity line beneath a forest plot.
Consistency across settings can reduce some concerns about effect modification
If an effect appears similar across populations or settings that differ on a proposed effect modifier, concern that the modifier dramatically changes the relative effect may decrease.
Core GRADE approaches indirectness by asking whether differences between the target and study PICO are likely to produce substantially different effects. It also notes that many population differences do not, in practice, generate important differences in relative effects.
Evidence of consistency across relevant variation can therefore contribute to the judgment that an observed finding may travel beyond one narrowly defined population.
But that inference should remain bounded by the contexts actually represented.
The studies still need to be credible
Consistency is not a substitute for study quality.
If several studies share the same bias, the same biased result can recur. Similar estimates may also arise because studies use the same flawed measurement, selective reporting practices, or implementation assumptions.
Five biased studies agreeing with one another do not become unbiased through teamwork.
Before treating cross-setting consistency as strengthening confidence, examine risk of bias, measurement validity, study design, precision, and whether the studies are sufficiently independent to provide meaningful replication.
Independence matters
Apparently separate studies may not constitute independent replications. They may analyze overlapping participants, use the same dataset, reproduce the same intervention team, or share common methodological constraints.
Similarly, multicenter studies provide valuable evidence across locations, but centers may still operate under a tightly standardized protocol that minimizes meaningful contextual variation.
Ask what has genuinely been replicated: the numerical analysis, the intervention under tightly controlled conditions, or the underlying effect across different environments?
Consistency can coexist with different absolute effects
Relative effects may remain reasonably consistent while absolute effects vary because baseline risks differ.
Suppose an intervention reduces relative risk by approximately 20% across several populations. In a setting where baseline risk is 50%, that corresponds to a much larger absolute reduction than in a setting where baseline risk is 5%.
The relative effect may therefore appear robust across contexts while the practical consequences differ substantially.
When evaluating consistency, specify which effect measure is consistent.
Contextual diversity can support, but not prove, generalizability
Evidence observed across multiple relevant contexts can support a broader inference than evidence observed under one narrow set of conditions. Yet no finite collection of settings establishes universal generalizability.
New settings may introduce a characteristic absent from all previous studies. An intervention consistently effective across hospitals with reliable electricity has not thereby demonstrated effectiveness where electricity is intermittent if electricity is essential to delivery.
That is why limits to generalization still need to be considered even when the evidence looks impressively consistent.
Different effects can be informative too
Consistency is not the only scientifically useful pattern. If an effect changes systematically when context changes, the variation may identify an important condition under which the intervention operates.
In that situation, forcing the studies toward a single average may sacrifice explanatory information. Different effects across settings can reveal a contextual mechanism when the pattern is credible and competing explanations have been considered.