03 · What You Need to Know
Country is often a label for several potentially important differences
A national border does not itself create methodological incomparability. Two studies conducted thousands of kilometers apart may examine remarkably similar populations under similar institutional conditions. Conversely, two studies conducted within the same country may operate in settings so different that treating them as equivalent would be difficult to justify.
The useful question is therefore not simply, “Were these studies conducted in different countries?” It is, “Which differences represented by those settings could matter for the question being synthesized?”
Start with the question each study actually answers
Before considering country differences, establish whether the studies are conceptually comparable. Compare their populations, exposures or interventions, comparison conditions, outcomes, designs, measurement approaches, and relevant time periods.
For intervention evidence, the population, intervention, comparator, and outcome framework provides one useful starting point. The GRADE approach similarly treats evidence as potentially indirect when the study population, intervention, comparator, or outcome differs meaningfully from the target question. The important issue is not difference in the abstract, but whether that difference creates a credible possibility that the magnitude or interpretation of the effect would change.
This means that a study conducted in another country is not automatically indirect evidence. Nor is evidence automatically direct simply because it comes from your own country.
Ask what the country difference actually represents
“Country” is usually too coarse to function as an explanation by itself. A difference between findings from two countries might reflect differences in health systems, educational structures, income distribution, language, infrastructure, regulation, cultural norms, access to technology, baseline risk, institutional capacity, implementation fidelity, or numerous other conditions.
For example, suppose several studies evaluate an online learning intervention. Calling one study “Philippine evidence” and another “Australian evidence” tells you relatively little about why their findings might differ. More informative questions concern internet access, device availability, class size, teacher preparation, curriculum, student characteristics, and how the intervention was implemented.
Country as location
The study happened within a particular national boundary, but that fact may have little substantive importance for the research question.
Country as context
The setting contains institutional, social, economic, cultural, epidemiological, or implementation conditions that could plausibly alter the phenomenon being studied.
Separate conceptual heterogeneity from statistical heterogeneity
Studies can differ conceptually even when their numerical estimates look similar. Likewise, estimates may differ numerically even though the studies address a sufficiently common underlying question.
Statistical heterogeneity concerns variation in observed effect estimates beyond what would be expected from sampling variation alone. Conceptual or clinical heterogeneity concerns differences in populations, interventions, outcomes, designs, or settings. The two are related, but they are not interchangeable.
A low statistical heterogeneity statistic does not prove that several national settings are substantively equivalent. With few studies, heterogeneity may also be estimated imprecisely. Context therefore needs to be considered from knowledge of the studies and the phenomenon, not merely from a statistic produced after pooling.
Look for plausible effect modifiers before combining results
A contextual characteristic becomes especially important when there is a credible reason to expect it to modify the relationship being studied.
Imagine an intervention whose effectiveness depends heavily on trained personnel. Results from countries with very different staffing capacity might legitimately differ because the intervention cannot be delivered in the same way. Similarly, the effectiveness of a digital service may depend on connectivity, device ownership, digital literacy, or institutional support.
The appropriate contextual variables depend on the research question. There is no universal checklist of national characteristics that must always match.
Watch Out
Do not use broad country categories as convenient explanations without establishing what they represent. Labels such as “developed,” “developing,” “Western,” or even broad income classifications can conceal substantial variation within groups. If context appears to explain differences, identify the contextual feature as precisely as the available evidence allows.
Consider whether the intervention itself changes across countries
An intervention with the same name may not be the same intervention in practice. Delivery intensity, staffing, implementation procedures, eligibility criteria, available resources, adherence, and the surrounding standard of care can all differ.
This is especially consequential when the comparator changes. “Usual care,” for example, may represent quite different services in two health systems. An intervention may therefore show a different incremental benefit not because its intrinsic effect changed, but because the alternative against which it was compared changed.
The same issue occurs outside health research. Business-as-usual teaching, conventional public services, workplace practices, and regulatory environments can vary substantially across countries.
Check whether outcomes mean the same thing across settings
Researchers may use identically named outcomes without measuring precisely the same construct. Translation, cultural adaptation, response styles, administrative definitions, diagnostic practices, grading systems, and measurement instruments can affect comparability.
A measure validated in one language or population should not automatically be assumed to function identically elsewhere. For qualitative synthesis, meanings attached to concepts such as autonomy, stigma, trust, family responsibility, or professional authority may also depend heavily on context.
Sometimes harmonization is defensible. Sometimes the difference itself belongs in the interpretation.
Do not confuse synthesis with statistical pooling
Evidence can be synthesized without forcing every result into a single pooled estimate. A review may organize studies around common questions, compare patterns across settings, examine contextual explanations, or use structured narrative synthesis when quantitative pooling would obscure important differences.
Meta-analysis is therefore one possible method within evidence synthesis, not the definition of synthesis itself.
If the studies are sufficiently comparable for quantitative synthesis, contextual heterogeneity should still inform interpretation. Where enough studies and suitable data exist, prespecified subgroup analysis or meta-regression may help investigate whether particular contextual characteristics are associated with variation in effects. Such analyses need caution, particularly when based on small numbers of studies or post hoc explanations.
Think separately about synthesis and applicability
Two questions are easy to collapse into one:
Can these studies be synthesized?
Are they sufficiently comparable to contribute to a coherent evidence synthesis?
Does this synthesis apply to my setting?
Is the resulting evidence sufficiently direct and relevant to the population, institutions, resources, and conditions where you want to use it?
A cross-country synthesis can be methodologically coherent while remaining uncertain in its applicability to a particular setting. Conversely, evidence from another country may sometimes be highly relevant when the characteristics that matter for the question closely resemble your target context.
This is why geographical proximity is a poor substitute for a substantive assessment of whether evidence from another population is relevant to yours.
Variation across countries can sometimes be informative
Researchers often approach heterogeneity as a problem to eliminate. Yet variation can contain substantive information. If an effect appears under very different contextual conditions, that pattern may contribute to confidence that the finding is not restricted to one narrow environment. This does not prove universal generalizability, but consistency across substantially different settings can be informative.
The opposite pattern may be equally valuable. If effects differ systematically with institutional capacity, resource availability, policy environment, or another plausible contextual factor, the heterogeneity may help reveal how context shapes the underlying mechanism.
In those circumstances, averaging away the difference could discard part of the finding.