01 · The Question
How Diverse Is the Evidence Behind Your Conclusion?
Your review contains 25 studies, which sounds reassuring. Then you examine them more closely. Fifteen come from one country. Nine involve the same research group. Several use the same national dataset. Most participants come from similar institutions.
You still have 25 studies, but perhaps not 25 independent tests of whether the finding travels across settings.
Concentration is not automatically a flaw. One country may legitimately dominate research on a locally specific policy, one institution may maintain a uniquely valuable cohort, and one research group may simply be especially productive. The problem arises when a concentrated evidence base is interpreted as though it represents populations, institutions, or contexts it has barely studied.
03 · What You Need to Know
A Large Literature Can Still Come From a Narrow Evidence Base
Study count does not tell you how broadly the evidence has been tested
Twenty studies conducted across twenty substantially different settings provide a different kind of information from twenty studies repeatedly conducted in one setting. That does not mean the first set is automatically better. It means the two evidence bases support different judgments about contextual replication and applicability.
When studies cluster around the same institutions, researchers, populations, or data systems, their apparent numerical abundance can obscure a narrower evidence base underneath.
Country can matter, but nationality is not itself the mechanism
Finding that most studies come from one country should prompt questions, not an automatic penalty. The relevant issue is whether features that vary across settings could affect the phenomenon or intervention.
Cochrane's guidance on applicability identifies contextual factors such as cultural and linguistic diversity, socioeconomic position, rural or urban setting, and characteristics of service delivery as potentially relevant to whether evidence transfers from one setting to another. It advises caution when generalizing across contexts.
The implication is not that evidence from one country is intrinsically weak. Rather, broad international claims require justification when the underlying evidence comes from a comparatively narrow set of contexts.
Internal validity
Whether a study's findings are credible for the participants and conditions actually studied.
Applicability or generalizability
Whether and to what extent those findings can reasonably inform other populations, settings, or circumstances.
A study can be methodologically rigorous yet provide indirect evidence for a substantially different target setting. Cochrane and GRADE address this issue through the concept of indirectness, which asks how closely the available evidence aligns with the population, intervention, comparator, and outcomes in the question of interest.
Institutional concentration can hide contextual dependencies
Research repeatedly conducted at one university, hospital system, laboratory, school network, or company may share characteristics that are not obvious from study titles. Recruitment practices, infrastructure, staff expertise, socioeconomic context, implementation support, institutional culture, or access to technology may influence the observed results.
This matters particularly for complex interventions. A program that succeeds in a highly resourced research-intensive institution may not perform identically when implemented with different personnel, infrastructure, incentives, or participant populations.
The appropriate response is not to dismiss single-institution research. It is to avoid silently converting evidence from a specific environment into a universal claim.
Repeated use of one dataset can create the appearance of replication
Several papers can analyze the same national survey, cohort, registry, administrative database, or institutional dataset. They may ask different questions and make legitimate contributions, but they do not necessarily provide independent replication across populations.
If two analyses use overlapping participants, there may also be a statistical dependency that requires attention. Before interpreting several papers as several independent confirmations, determine whether you need to account for the same participants appearing more than once.
Research-group concentration creates another kind of dependence
Studies from the same research team may be statistically independent while still sharing conceptual and methodological features. Researchers often reuse instruments, analytic conventions, recruitment networks, implementation procedures, theoretical assumptions, or operational definitions across projects.
Replication by the originating team and replication by independent teams therefore answer somewhat different questions. Consistent findings from one productive group may show that an effect is reproducible within that research program, but independent replication can provide additional information about whether the finding survives different investigators, procedures, and settings.
This should not be turned into suspicion based on authorship alone. Repeated authorship is a signal to inspect the structure of the evidence, not evidence that the research is biased.
Search systems can contribute to geographic concentration
Your final evidence base partly reflects where you searched. Database coverage differs across disciplines, journals, languages, and regions. Restricting a search to familiar international databases can make some research ecosystems easier to discover than others.
If geographic breadth matters to the question, consider whether your choice of databases provides adequate coverage of relevant regional literature. Vocabulary can matter too, particularly when terminology differs across disciplines or linguistic contexts.
Language restrictions can narrow the evidence base
Language eligibility and database coverage are separate but related issues. A search may locate studies in several languages only for the review to exclude most of them because of a language restriction. Conversely, a supposedly language-inclusive review cannot include literature its search sources never expose.
Restrictions may sometimes be necessary because of resources or review scope. They should nevertheless be reported transparently, particularly when they could affect the geographic or cultural composition of the evidence.
Map concentration before deciding whether it is a problem
A simple evidence map can make concentration visible. Record where studies were conducted, which institutions and research groups recur, which datasets are reused, and how participants are distributed across relevant contexts.
| Dimension |
What to inspect |
Why it may matter |
| Country or region |
Distribution of studies and participants |
Social, cultural, policy, economic, or service contexts may differ |
| Institution |
Repeated universities, hospitals, schools, laboratories, or organizations |
Institution-specific resources and practices may influence results |
| Research group |
Recurring investigators and collaborative networks |
Studies may share methods, assumptions, instruments, or implementation practices |
| Dataset or cohort |
Repeated named cohorts, surveys, registries, or administrative datasets |
Several papers may not represent independent populations |
| Population |
Age, socioeconomic characteristics, language, ethnicity where relevant and appropriately reported, urbanity, or other contextual features |
The evidence may not directly represent the target population |
| Setting |
Clinical, educational, workplace, community, laboratory, online, or other environments |
Effects can depend on implementation conditions and context |
Do not manufacture diversity by giving weak evidence extra weight
Suppose most high-quality studies come from one region while a few methodologically weak studies come from elsewhere. Geographic diversity does not require treating those weaker studies as equally credible merely to create balance.
Contextual breadth and methodological strength are separate dimensions. You should examine both. A geographically narrow but rigorous evidence base may warrant cautious generalization, while a geographically broad collection of severely biased studies does not become strong evidence simply because the map looks impressive.
The solution is to give stronger and weaker evidence appropriate weight while separately discussing limitations in applicability.
Sometimes concentration accurately describes the research field
A concentrated evidence base may persist even after a broad, multilingual, multi-database search. If so, the concentration itself is a finding about the state of the literature.
Do not imply that evidence exists elsewhere merely because you wish the literature were more diverse. Instead, distinguish a biased search process from a genuine evidence gap. If most available studies really come from a small number of contexts, say so and calibrate the conclusion accordingly.
04 · A Practical Example
When 18 Studies Represent Fewer Contexts Than the Number Suggests
Hypothetical Example
A review of AI-supported feedback in higher education
A researcher identifies 18 studies evaluating AI-supported feedback for university students. At first glance, the literature seems reasonably substantial.
Map the studies
The researcher records country, institution, participant population, research group, and data source for every study.
Identify concentration
Ten studies come from universities in one country. Six of those involve the same research group, and four use students from the same institution.
Check independence
The researcher determines that two papers use partially overlapping cohorts and avoids treating them as fully independent replications.
Examine context
The remaining studies come from several settings, but few examine institutions with substantially different technological infrastructure or student populations.
Calibrate the conclusion
Instead of claiming that AI-supported feedback is effective across higher education generally, the review reports that favorable findings have been observed predominantly in a limited set of institutional and geographic contexts, with broader transferability remaining less certain.
The number of included studies has not changed. What changed is the claim the evidence can reasonably support.
07 · A Quick Checklist
Before You Generalize Beyond the Settings You Reviewed
Map the provenance of your evidence:
I recorded the countries or regions in which the included studies were conducted when geography is relevant to applicability.
I checked whether a small number of institutions or research groups produced a large share of the evidence.
I identified repeated cohorts, surveys, registries, or other datasets rather than assuming each paper represents an independent population.
I considered whether populations and settings in the evidence match those to which I want to apply the conclusion.
I considered whether database or language choices may have made evidence from some regions easier to find than others.
I kept methodological quality separate from geographic or institutional diversity.
I distinguished a narrow search from a genuinely narrow available evidence base.
The geographic and contextual breadth of my conclusion does not exceed what the evidence can reasonably support.