Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Synthesize Evidence Across Clinical and Nonclinical Populations?

Clinical and nonclinical populations can sometimes contribute to the same synthesis, but only when they address a genuinely common question. Diagnosis alone should neither force separation nor justify pooling.

571
Clinical and Nonclinical Evidence Synthesis Guide 571 of 899
01 · The Question

Can findings from patients and nonclinical populations be combined?

Imagine that your review concerns anxiety, cognitive performance, physical activity, sleep, digital behavior, or another phenomenon studied both in diagnosed patients and in people recruited from schools, workplaces, universities, or the general community.

The measures may look similar. The intervention may even have the same name. Yet the populations differ in ways that could matter.

Should clinical and nonclinical samples contribute to one synthesis, or would combining them create an average that describes neither population particularly well?

02 · The Short Answer

Population labels alone should not determine the synthesis

In Brief

Clinical and nonclinical populations can be synthesized when they address a sufficiently common question and differences in diagnosis, severity, baseline risk, treatment context, measurement, or intervention delivery do not make their findings substantively incomparable.

When those characteristics could plausibly modify the finding, preserve the distinction and investigate it rather than assuming one overall effect applies equally to both populations.

03 · What You Need to Know

The important difference is often not simply “clinical versus nonclinical”

A clinical sample is generally recruited because participants meet a health-related criterion, receive services in a clinical setting, or have a diagnosed or otherwise defined condition. A nonclinical sample may be recruited without that criterion from a community, educational institution, workplace, or general population.

Those labels are useful descriptions, but they can conceal considerable heterogeneity within each category.

A mildly symptomatic outpatient sample and a severely ill inpatient sample are both “clinical.” A community sample can contain people with substantial symptoms who have never received a diagnosis. Classification by recruitment source therefore does not necessarily create two biologically or behaviorally homogeneous populations.

Ask what clinical status represents for your particular question

The key issue is whether clinical status corresponds to characteristics capable of changing the finding.

Those characteristics might include baseline symptom severity, disease stage, comorbidity, medication use, previous treatment, referral pathways, functional impairment, baseline risk, motivation to participate, access to professional support, or the setting in which an intervention is delivered.

If these factors influence the outcome or intervention effect, clinical status may be an important marker of heterogeneity. If they do not, separating studies simply because one recruited through a clinic may be unnecessary.

Population label A descriptive classification such as clinical, patient, community, student, or general-population sample.
Effect-relevant characteristic A feature such as severity, baseline risk, comorbidity, treatment exposure, or service context that could plausibly alter the finding.

Baseline severity can change what an effect means

Suppose an intervention reduces symptom scores by a similar relative amount in clinical and nonclinical samples. Participants entering with severe symptoms may nevertheless experience a different absolute change or different clinical implications than participants beginning near the healthy end of the scale.

Floor and ceiling effects may also matter. A preventive intervention evaluated among people with minimal symptoms has less room to produce symptom reduction than the same intervention evaluated among participants selected for elevated symptoms.

Consequently, an apparent population difference may sometimes reflect different starting points rather than a fundamentally different mechanism.

Clinical status may be prognostic without modifying an intervention effect

A characteristic that strongly predicts outcomes does not necessarily change how an intervention works. Cochrane distinguishes prognostic factors from effect modifiers for precisely this reason. A characteristic is useful for subgroup analysis when there is reason to believe it modifies the intervention effect, not merely because it predicts the outcome.

Clinical participants may have worse outcomes overall while experiencing approximately the same relative intervention effect as nonclinical participants.

That distinction should influence both analysis and interpretation.

The intervention may not really be the same across settings

A named intervention can change substantially when moved between a clinic and a community setting.

Clinical delivery may involve trained specialists, diagnostic assessment, individualized monitoring, concurrent medication, referral options, or structured follow-up. A superficially similar program delivered through a school or smartphone app may have different intensity, adherence, safeguards, or supporting resources.

If intervention delivery changes with population setting, a clinical versus nonclinical contrast becomes entangled with an intervention contrast.

The comparator may change too. “Usual care” in a clinical trial may include services unavailable to a community sample. An intervention's incremental effect can therefore differ because the alternative against which it is evaluated differs.

Check whether the outcome is genuinely comparable

Clinical and nonclinical studies may use the same questionnaire but interpret its scores differently. Others may use diagnostic remission in clinical samples and continuous symptom scores in community samples.

Before combining outcomes, determine whether they represent sufficiently comparable constructs and whether the chosen effect measure preserves a meaningful interpretation.

Standardizing different scales can sometimes permit quantitative synthesis, but statistical standardization does not establish conceptual equivalence. Two measures can be converted to the same mathematical scale while still assessing meaningfully different outcomes.

Clinical and nonclinical populations may overlap

The boundary between these populations is often less clean than the terminology suggests.

Community samples may contain diagnosed individuals. Clinical samples may include participants across a broad severity range. Diagnostic criteria can differ between studies, and thresholds can change over time or across classification systems.

For some questions, symptom severity or baseline risk may therefore be a more informative synthesis variable than the binary clinical/nonclinical label.

Investigate subgroup differences carefully

If clinical status is a plausible effect modifier, subgroup analysis or meta-regression may help investigate heterogeneity. However, Cochrane emphasizes that subgroup analyses are observational and vulnerable to false-positive and false-negative findings, especially when many analyses are conducted or few studies are available.

Between-study comparisons deserve particular caution. If all clinical studies were conducted in hospitals and all nonclinical studies online, any apparent population difference is confounded with delivery setting. If clinical studies are older, smaller, or at different risk of bias, those characteristics may also contribute.

Watch Out

Do not conclude that an intervention “works in clinical populations but not nonclinical populations” merely because one subgroup reaches statistical significance and the other does not. Compare the subgroup effects directly and consider whether population status is confounded with other study characteristics.

Sometimes the populations should remain separate

Separate synthesis becomes more compelling when clinical status changes the underlying question rather than merely describing the sample.

A treatment for diagnosed disease and a preventive intervention among asymptomatic people may pursue different goals. Outcomes, acceptable harms, baseline risks, delivery pathways, and meaningful effect sizes can all differ. Combining them might produce a mathematically valid average with weak substantive meaning.

Cochrane distinguishes clinical diversity in populations, interventions, and outcomes from methodological and statistical heterogeneity, and cautions that diversity must be considered regardless of the synthesis method chosen.

Sometimes combining them answers the more interesting question

There are also circumstances in which a phenomenon genuinely spans clinical boundaries. If studies examine a common mechanism across a severity continuum and the intervention, comparator, outcome, and causal process remain sufficiently comparable, a broader synthesis may be defensible.

In that situation, clinical status can be preserved as a potential moderator without requiring two completely disconnected reviews.

The broader principle is the same as when deciding whether evidence across age groups can be synthesized: population differences matter when they alter the question or the interpretation, not merely because demographic or diagnostic labels differ.

Do not confuse synthesis with generalization

You may legitimately synthesize clinical and nonclinical evidence to answer a broad research question while remaining uncertain whether the resulting estimate applies to one specific target population.

Applicability should therefore be assessed separately. If the evidence base is dominated by community participants, a clinician should not automatically assume that the pooled finding transfers unchanged to patients with severe disease. Conversely, evidence derived entirely from specialist clinics may not directly describe people with mild symptoms in the general population.

When the target population differs substantially from those represented, ask whether evidence from another population is directly relevant rather than treating inclusion in the same synthesis as proof of generalizability.

04 · A Practical Example

When one intervention crosses the clinical boundary

Hypothetical Example

A digital sleep intervention in clinical and community samples

Imagine a hypothetical systematic review of a digital intervention intended to improve sleep. Some studies recruit patients diagnosed with insomnia from clinics. Others recruit adults from the community who report sleep difficulties but are not required to have a diagnosis.

Common ground The studies evaluate substantially similar intervention content and use comparable measures of sleep symptoms.
Important difference Clinical participants generally begin with more severe symptoms and some receive concurrent treatment, while community participants tend to have milder difficulties.
Initial synthesis The reviewers judge that the studies address a sufficiently common overarching question to be synthesized, while preserving clinical status and baseline severity during extraction.
Investigating variation Clinical studies appear to show larger absolute symptom reductions, but relative standardized effects are less clearly different between populations.
Interpretation The reviewers do not immediately conclude that the intervention intrinsically works better in clinical populations. Greater baseline severity, concurrent care, intervention support, and other study differences remain plausible explanations.

The useful finding is not necessarily a binary verdict about clinical versus nonclinical populations. The synthesis may instead reveal how baseline severity, treatment context, and implementation shape the observed effect.

05 · What Researchers Often Get Wrong

Common mistakes when combining clinical and nonclinical evidence

Misconception

Clinical and nonclinical samples should never be pooled

The label alone does not determine comparability. If the studies address a genuinely common question and effect-relevant characteristics remain sufficiently comparable, a broader synthesis may be appropriate.

Misconception

A community sample contains only healthy participants

Community recruitment does not guarantee absence of symptoms, diagnoses, medication use, or previous treatment. Examine actual eligibility criteria and participant characteristics rather than inferring them from recruitment setting.

Misconception

Clinical status is the same thing as severity

It is not. Clinical populations can contain broad severity ranges, while nonclinical samples may include highly symptomatic participants. When severity is the plausible effect modifier, analyze it as directly as the available evidence permits.

Misconception

Using the same questionnaire makes the populations comparable

Measurement similarity is only one component of comparability. Baseline risk, intervention delivery, comparator conditions, comorbidity, treatment context, and the meaning of the outcome may still differ.

Misconception

A pooled effect automatically applies to both populations

An overall synthesis can be coherent while applicability remains uncertain for particular populations. Representation and credible effect modification still need to be considered.

06 · What This Means for You

Decide whether clinical status changes the question before splitting the evidence

Start by defining the phenomenon, intervention or exposure, comparator, and outcome. Then ask what the clinical/nonclinical distinction represents in the included studies.

A simple decision framework

If clinical and nonclinical studies investigate essentially the same phenomenon under sufficiently comparable conditions
A combined synthesis may be appropriate while preserving population characteristics for interpretation.
If diagnosis or clinical status corresponds to plausible effect modifiers such as severity or baseline risk
Prespecify relevant subgroup or moderator analyses where the available evidence can support them.
If intervention delivery or comparator conditions systematically differ between clinical and community settings
Avoid attributing observed differences to population status alone.
If the populations represent fundamentally different research questions, such as treatment versus prevention
Separate syntheses will often provide a clearer and more interpretable answer.
If your target population is poorly represented
Explicitly consider indirectness and avoid assuming that the overall estimate generalizes unchanged.

Resource conditions may further complicate the distinction. A specialist clinical population treated in a highly resourced center may differ from a community population not only in diagnosis but also in staffing, monitoring, technology, and service availability. When those factors are consequential, resource setting itself may need explicit treatment in the synthesis.

If meaningful differences remain and the evidence cannot support a credible common interpretation, saying that the evidence does not generalize directly is preferable to presenting a broad average as universally applicable.

07 · A Quick Checklist

Before combining clinical and nonclinical populations, check:

Before synthesizing these populations together, check:
Do the studies answer the same substantive research question?
How does each study actually define its clinical or nonclinical population?
Do baseline severity, risk, comorbidity, or previous treatment differ in ways that could matter?
Is the intervention genuinely comparable across clinical and nonclinical settings?
Are comparator conditions sufficiently similar?
Do the outcomes measure sufficiently comparable constructs in both populations?
Could apparent population differences actually reflect setting, implementation, study design, or risk of bias?
Would one pooled estimate conceal a clinically or substantively important difference?
Is your intended target population adequately represented in the evidence?
08 · Frequently Asked Questions

Questions about clinical and nonclinical evidence

Can clinical and community samples be included in the same systematic review?

Yes, when both contribute meaningfully to the review question. Inclusion in the same review does not necessarily mean that every study should be combined in the same quantitative synthesis.

Should clinical status automatically be a subgroup in meta-analysis?

No. Subgroup analyses should have a substantive rationale and enough evidence to be informative. If severity or treatment context is the more plausible modifier, those characteristics may be more useful than a broad clinical/nonclinical label.

Can I combine diagnosed participants with people who have elevated symptoms but no diagnosis?

Potentially, depending on the research question. Examine whether diagnosis marks a meaningful difference in severity, underlying condition, treatment context, intervention delivery, or outcome interpretation rather than treating diagnostic status as an automatic boundary.

What if clinical studies show larger effects?

Investigate whether the difference is credible and what might explain it. Baseline severity, concurrent treatment, intervention intensity, comparator conditions, setting, study methods, or genuine effect modification are all possibilities.

Can standardized mean differences solve population comparability problems?

No. Standardization can place outcomes measured on different numerical scales onto a common effect-size metric, but it cannot make conceptually different populations, interventions, or outcomes equivalent.

What if the evidence is almost entirely clinical but my target population is nonclinical?

Assess whether characteristics distinguishing the target population could plausibly alter the finding. If so, applicability may be uncertain even when the clinical evidence itself is internally convincing.

09 · The Bottom Line

Clinical status matters when it changes what the evidence means

The Bottom Line

Clinical and nonclinical populations can be synthesized when they contribute to a genuinely common question, but they should remain distinct when diagnosis, severity, baseline risk, treatment context, intervention delivery, or outcome meaning changes the interpretation of the evidence.

Do not let a population label do the analytical work for you. Identify what actually differs between the populations, determine whether those characteristics could modify the finding, and separate synthesis from the later question of whether the evidence applies to your target population.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes