01 · The Question
Can findings from patients and nonclinical populations be combined?
Imagine that your review concerns anxiety, cognitive performance, physical activity, sleep, digital behavior, or another phenomenon studied both in diagnosed patients and in people recruited from schools, workplaces, universities, or the general community.
The measures may look similar. The intervention may even have the same name. Yet the populations differ in ways that could matter.
Should clinical and nonclinical samples contribute to one synthesis, or would combining them create an average that describes neither population particularly well?
03 · What You Need to Know
The important difference is often not simply “clinical versus nonclinical”
A clinical sample is generally recruited because participants meet a health-related criterion, receive services in a clinical setting, or have a diagnosed or otherwise defined condition. A nonclinical sample may be recruited without that criterion from a community, educational institution, workplace, or general population.
Those labels are useful descriptions, but they can conceal considerable heterogeneity within each category.
A mildly symptomatic outpatient sample and a severely ill inpatient sample are both “clinical.” A community sample can contain people with substantial symptoms who have never received a diagnosis. Classification by recruitment source therefore does not necessarily create two biologically or behaviorally homogeneous populations.
Ask what clinical status represents for your particular question
The key issue is whether clinical status corresponds to characteristics capable of changing the finding.
Those characteristics might include baseline symptom severity, disease stage, comorbidity, medication use, previous treatment, referral pathways, functional impairment, baseline risk, motivation to participate, access to professional support, or the setting in which an intervention is delivered.
If these factors influence the outcome or intervention effect, clinical status may be an important marker of heterogeneity. If they do not, separating studies simply because one recruited through a clinic may be unnecessary.
Population label
A descriptive classification such as clinical, patient, community, student, or general-population sample.
Effect-relevant characteristic
A feature such as severity, baseline risk, comorbidity, treatment exposure, or service context that could plausibly alter the finding.
Baseline severity can change what an effect means
Suppose an intervention reduces symptom scores by a similar relative amount in clinical and nonclinical samples. Participants entering with severe symptoms may nevertheless experience a different absolute change or different clinical implications than participants beginning near the healthy end of the scale.
Floor and ceiling effects may also matter. A preventive intervention evaluated among people with minimal symptoms has less room to produce symptom reduction than the same intervention evaluated among participants selected for elevated symptoms.
Consequently, an apparent population difference may sometimes reflect different starting points rather than a fundamentally different mechanism.
Clinical status may be prognostic without modifying an intervention effect
A characteristic that strongly predicts outcomes does not necessarily change how an intervention works. Cochrane distinguishes prognostic factors from effect modifiers for precisely this reason. A characteristic is useful for subgroup analysis when there is reason to believe it modifies the intervention effect, not merely because it predicts the outcome.
Clinical participants may have worse outcomes overall while experiencing approximately the same relative intervention effect as nonclinical participants.
That distinction should influence both analysis and interpretation.
The intervention may not really be the same across settings
A named intervention can change substantially when moved between a clinic and a community setting.
Clinical delivery may involve trained specialists, diagnostic assessment, individualized monitoring, concurrent medication, referral options, or structured follow-up. A superficially similar program delivered through a school or smartphone app may have different intensity, adherence, safeguards, or supporting resources.
If intervention delivery changes with population setting, a clinical versus nonclinical contrast becomes entangled with an intervention contrast.
The comparator may change too. “Usual care” in a clinical trial may include services unavailable to a community sample. An intervention's incremental effect can therefore differ because the alternative against which it is evaluated differs.
Check whether the outcome is genuinely comparable
Clinical and nonclinical studies may use the same questionnaire but interpret its scores differently. Others may use diagnostic remission in clinical samples and continuous symptom scores in community samples.
Before combining outcomes, determine whether they represent sufficiently comparable constructs and whether the chosen effect measure preserves a meaningful interpretation.
Standardizing different scales can sometimes permit quantitative synthesis, but statistical standardization does not establish conceptual equivalence. Two measures can be converted to the same mathematical scale while still assessing meaningfully different outcomes.
Clinical and nonclinical populations may overlap
The boundary between these populations is often less clean than the terminology suggests.
Community samples may contain diagnosed individuals. Clinical samples may include participants across a broad severity range. Diagnostic criteria can differ between studies, and thresholds can change over time or across classification systems.
For some questions, symptom severity or baseline risk may therefore be a more informative synthesis variable than the binary clinical/nonclinical label.
Investigate subgroup differences carefully
If clinical status is a plausible effect modifier, subgroup analysis or meta-regression may help investigate heterogeneity. However, Cochrane emphasizes that subgroup analyses are observational and vulnerable to false-positive and false-negative findings, especially when many analyses are conducted or few studies are available.
Between-study comparisons deserve particular caution. If all clinical studies were conducted in hospitals and all nonclinical studies online, any apparent population difference is confounded with delivery setting. If clinical studies are older, smaller, or at different risk of bias, those characteristics may also contribute.
Watch Out
Do not conclude that an intervention “works in clinical populations but not nonclinical populations” merely because one subgroup reaches statistical significance and the other does not. Compare the subgroup effects directly and consider whether population status is confounded with other study characteristics.
Sometimes the populations should remain separate
Separate synthesis becomes more compelling when clinical status changes the underlying question rather than merely describing the sample.
A treatment for diagnosed disease and a preventive intervention among asymptomatic people may pursue different goals. Outcomes, acceptable harms, baseline risks, delivery pathways, and meaningful effect sizes can all differ. Combining them might produce a mathematically valid average with weak substantive meaning.
Cochrane distinguishes clinical diversity in populations, interventions, and outcomes from methodological and statistical heterogeneity, and cautions that diversity must be considered regardless of the synthesis method chosen.
Sometimes combining them answers the more interesting question
There are also circumstances in which a phenomenon genuinely spans clinical boundaries. If studies examine a common mechanism across a severity continuum and the intervention, comparator, outcome, and causal process remain sufficiently comparable, a broader synthesis may be defensible.
In that situation, clinical status can be preserved as a potential moderator without requiring two completely disconnected reviews.
The broader principle is the same as when deciding whether evidence across age groups can be synthesized: population differences matter when they alter the question or the interpretation, not merely because demographic or diagnostic labels differ.
Do not confuse synthesis with generalization
You may legitimately synthesize clinical and nonclinical evidence to answer a broad research question while remaining uncertain whether the resulting estimate applies to one specific target population.
Applicability should therefore be assessed separately. If the evidence base is dominated by community participants, a clinician should not automatically assume that the pooled finding transfers unchanged to patients with severe disease. Conversely, evidence derived entirely from specialist clinics may not directly describe people with mild symptoms in the general population.
When the target population differs substantially from those represented, ask whether evidence from another population is directly relevant rather than treating inclusion in the same synthesis as proof of generalizability.