03 · What You Need to Know
An Institution Can Be More Than a Location
Institutions select both people and conditions
A university does not contain a random cross-section of all students. A hospital does not treat a random cross-section of all patients. Institutions admit, recruit, refer, hire, serve, and retain particular populations.
Institutional characteristics can therefore shape the study sample before researchers recruit a single participant. A tertiary referral hospital may see more complicated cases. A selective university may enroll students with particular academic backgrounds. A specialist school may serve a population unlike that of ordinary schools.
This means that evaluating a single-institution study begins with whether the sample fits the population behind the authors' claims.
The site can also change the intervention itself
Institutional context affects more than who participates. Staffing, expertise, infrastructure, organizational culture, leadership, workload, resources, technology, routines, and implementation support can alter how an intervention is delivered.
A complex program implemented by its developers at their own institution may receive unusually intensive training, monitoring, troubleshooting, and enthusiasm. Another institution adopting the same nominal intervention may not reproduce those conditions.
Thus, the external-validity question sometimes concerns whether the treatment effect generalizes and sometimes whether the intervention as implemented in the original study can even be reproduced elsewhere.
Separate participant effects from site effects
Suppose an intervention performs exceptionally well at one hospital. At least two broad explanations deserve consideration. The hospital may treat patients who differ from those elsewhere, or characteristics of the hospital itself may improve treatment delivery and outcomes. Both can operate simultaneously.
Participant differences
The people entering the institution differ in prognosis, demographics, experience, resources, needs, or other relevant characteristics.
Institutional differences
The setting differs in expertise, staffing, resources, practices, implementation, policies, organizational structure, or other contextual features that influence the result.
These mechanisms have different implications. Statistical adjustment for participant characteristics may not remove an effect produced by institutional practices.
Single-site research cannot directly estimate between-site heterogeneity
A fundamental limitation of observing one institution is that there is no variation between institutions within the study. Researchers can estimate variation among individuals at that site, but they cannot directly determine from those data whether the result would be larger, smaller, or absent at another institution.
This does not prove that the finding is site-specific. It means that between-site stability has not been empirically tested by that study.
Multisite designs can provide evidence about variation across institutions, especially when participating sites differ meaningfully in populations and practice environments. However, merely adding sites does not guarantee broad generalizability if all participating institutions are themselves unusually similar or selectively chosen.
One institution may be exactly the right population
Sometimes the research question is deliberately local. A university may evaluate its own advising system to decide whether to continue it. A hospital may investigate waiting times in its emergency department. A school may study the implementation of a new local curriculum.
In these cases, generalization to other institutions may not be necessary for the primary purpose. The institution is not an inconvenient sample of a larger population; it is the population or setting of direct interest.
Problems arise when a local evaluation is subsequently written as though it established a general law about universities, hospitals, schools, or organizations.
Institutional similarity should be defined substantively
Authors sometimes suggest that findings should generalize to “similar institutions.” That phrase needs content. Similar in what respect?
For an educational intervention, relevant similarities might include student preparation, class size, curriculum, instructor workload, technological infrastructure, and assessment practices. For a clinical intervention, case mix, treatment expertise, staffing, referral patterns, baseline care, and available equipment may matter.
Two institutions can belong to the same formal category while differing on the characteristics that actually modify the result.
Prestigious or specialist institutions may be particularly unusual
Research-intensive universities, tertiary hospitals, specialist centers, and demonstration sites are often attractive places to conduct research because they have expertise and infrastructure. Those same advantages can make them atypical.
A large South Korean comparison of hepatocellular carcinoma cohorts, for example, found differences in survival outcomes between a high-volume single-center cohort and a nationwide multicenter cohort even after adjustment, illustrating how center volume, patient composition, and treatment strategies can affect the apparent portability of single-center results.
The lesson is not that research from specialist institutions should be distrusted. Rather, ask whether the institution's exceptional characteristics are part of the causal environment producing the result.
Site selection matters in multisite studies too
Multisite research is often described as more generalizable because it includes variation across settings. That can be true, but only if the sites provide useful diversity relative to the target settings.
If researchers select only enthusiastic, well-resourced institutions with experienced staff, the study may still provide weak evidence about ordinary sites with fewer resources. Site-level convenience sampling can create an external-validity problem even when participant numbers are large.
Research on pragmatic and community-based trials has therefore emphasized that representativeness can concern study sites as well as individual participants.
Replication at another institution is especially informative
When an effect depends heavily on local practices, replication at an independent institution tests something the original study cannot: whether the finding survives a change in organizational context.
Replication becomes particularly persuasive when the new site differs on plausible effect modifiers yet produces a compatible result. Conversely, a failure to reproduce the result does not automatically prove the original study was wrong. It may reveal genuine contextual heterogeneity worth explaining.
Prediction models need particularly careful external validation
Models developed and tested within one institution can exploit patterns specific to that institution's patient population, coding practices, equipment, workflows, or data-generating systems. Internal cross-validation does not expose the model to those external differences.
For prediction research, testing on genuinely independent institutions can therefore provide critical evidence about transportability. Recent multicenter research has shown that model performance can vary substantially across hospitals, reinforcing the need to validate models in populations and institutions resembling their intended deployment settings.
Do not confuse a large single-site sample with many settings
A hospital database containing 100,000 patients provides enormous information about patients treated within that system. It still represents one institutional environment.
Increasing the number of individuals improves precision and may capture substantial patient heterogeneity. It does not create institutional heterogeneity. This parallels the broader lesson that sample size cannot substitute for the dimensions of variation required by the research question.
Country and institution are separate levels of generalization
A single institution in one country creates at least two potential inferential steps: from the institution to other institutions within the country, and from that national context to settings abroad.
Researchers should not skip the first step. Before asking whether a result travels internationally, consider whether it has been shown to travel beyond the institution where it was produced. The broader question of generalizing findings across countries then introduces additional contextual dimensions.
Watch Out
Do not assume that “single institution” automatically means poor research or that “multicenter” automatically means generalizable research. What matters is whether the participating settings capture the variation relevant to the intended inference.