01 · The Question
Does evidence that something can work tell you how well it will work in routine practice?
An intervention performs well in a carefully controlled study. Participants are selected according to strict criteria. Providers receive extensive training. Adherence is closely monitored. Implementation problems are minimized.
Then the intervention moves into ordinary practice, where participants are more diverse, providers have competing demands, adherence varies, resources differ, and implementation is considerably less tidy.
Should evidence from these two settings be synthesized as though it answers exactly the same question?
Not necessarily. Evidence about efficacy and evidence about real-world effectiveness concern closely related but distinguishable questions. The first asks whether an intervention produces the intended effect under conditions designed to demonstrate that effect clearly. The second asks what happens when the intervention is used under conditions closer to routine practice.
02 · The Short Answer
Ask what conditions produced the observed effect
In Brief
Distinguish efficacy from real-world effectiveness by examining not only whether an intervention changes an outcome, but also the participants, settings, implementation conditions, comparators, adherence, follow-up, and delivery systems under which that effect was estimated.
Efficacy and effectiveness are better understood as positions along a continuum than as two perfectly separate categories. A study can be pragmatic in some respects and highly controlled in others, so synthesis should examine the actual design features that affect applicability rather than relying only on labels such as “efficacy trial” or “pragmatic trial.”
03 · What You Need to Know
The question is not only whether the intervention works, but under what conditions
Efficacy asks whether an intervention can produce an effect under specified conditions
Efficacy studies are generally designed to determine whether an intervention produces a desired effect under conditions that permit the intervention's effect to be estimated clearly. Those conditions may involve carefully selected participants, standardized delivery, close monitoring, high adherence, specialist providers, or other procedures that reduce variation unrelated to the intervention itself.
Such control can be scientifically valuable. If researchers are trying to determine whether an intervention has the capacity to produce an effect, allowing uncontrolled implementation failure to dominate the experiment may obscure the question they are trying to answer.
Effectiveness asks what happens under conditions closer to ordinary practice
Effectiveness evidence is more concerned with performance in the settings where the intervention is intended to be used. Pragmatic trials are specifically designed to inform decisions about practice, with attention to the applicability of results to real-world conditions.
That may mean broader participant eligibility, ordinary practitioners, routine service settings, realistic adherence, existing resources, usual-care comparators, and outcomes that matter to people making practical decisions.
Efficacy
Whether an intervention produces the intended effect under conditions selected to evaluate that effect clearly.
Effectiveness
What effect the intervention produces when delivered under conditions closer to the practice setting in which it is intended to be used.
The distinction should not be reduced to “laboratory versus real world.” Many efficacy trials occur in ordinary clinical, educational, organizational, or community settings, while pragmatic studies may still retain substantial experimental control.
Explanatory and pragmatic are not all-or-nothing labels
A particularly important mistake is to imagine two boxes: perfectly explanatory trials on one side and perfectly pragmatic trials on the other.
The CONSORT extension for pragmatic trials emphasizes that pragmatism is not an all-or-none property. A trial may have broad eligibility criteria while using an unusually intensive intervention protocol. Another may use ordinary practitioners and routine delivery but measure a short-term surrogate outcome that is less relevant to everyday decision-making.
For evidence synthesis, inspect the features of the study rather than trusting the label authors give it.
Feature
More efficacy-oriented condition
More effectiveness-oriented condition
Participants
Highly selected according to restrictive criteria
Broadly representative of intended users
Setting
Specialized or unusually well-resourced
Routine service or practice environment
Providers
Specially selected or extensively trained
Typical practitioners who would ordinarily deliver the intervention
Intervention delivery
Tightly standardized and closely monitored
Allows realistic variation in routine implementation
Adherence
Actively optimized
Reflects adherence likely to occur in practice
Comparator
May be placebo or tightly specified control
Often an existing practice alternative
Outcomes
May emphasize biological, behavioral, or intermediate responses
Often emphasizes outcomes important to practical decision-makers
The same intervention can have different effects under different implementation conditions
Suppose an intervention requires weekly specialist supervision. In an efficacy study, researchers ensure that almost every session occurs and that providers follow the protocol closely. In routine practice, staffing shortages may mean that participants receive only half of the intended sessions.
A smaller real-world effect would not necessarily contradict the efficacy trial. The intervention as implemented is no longer identical in every practically relevant respect to the intervention under tightly supported conditions.
This is where evidence about why an intervention works and how it is implemented becomes useful. Implementation evidence may help explain why effects change when the intervention moves across settings.
Eligibility criteria affect who the evidence applies to
A narrowly selected study population can strengthen certain aspects of an efficacy evaluation while limiting direct applicability to people who would receive the intervention in practice.
For example, a trial might exclude people with comorbidities, low baseline adherence, language barriers, or other characteristics common in the eventual target population. A positive result establishes an effect for the population and conditions actually studied more directly than it establishes the same effect for everyone who might later receive the intervention.
That does not make the trial uninformative. It means that transport from the study population to a broader population requires an additional judgment.
The comparator also affects the efficacy-effectiveness question
An intervention that performs well against placebo may provide a different practical answer from an intervention compared with a strong existing treatment. Real-world decision-makers often want to know what happens if the new option replaces or supplements what they already do.
Consequently, differences in comparison groups should remain visible when efficacy and effectiveness evidence are interpreted together.
Real-world evidence is not automatically better evidence
The phrase “real world” can sound inherently superior. It is not.
Greater resemblance to routine practice can improve applicability to a particular decision, but less control can also introduce methodological complications. An observational analysis of routine data may closely represent everyday practice while remaining vulnerable to confounding, measurement problems, missing data, or selection bias.
Conversely, a carefully conducted randomized efficacy trial may provide stronger evidence that the intervention caused an observed outcome difference while providing less direct information about how large that effect will be under routine implementation.
Internal validity and applicability are different considerations. One should not be sacrificed rhetorically to praise the other.
Do not assume an efficacy-effectiveness gap before examining the evidence
It is common to expect intervention effects to become smaller outside controlled studies. Sometimes they do. But that is not a universal law.
Routine settings can differ from efficacy studies in many directions. Real-world populations may have higher baseline risk and therefore greater potential absolute benefit. Providers may adapt an intervention productively. Additional services may complement it. Alternatively, poorer adherence or fewer resources may reduce its effect.
The direction and magnitude of any difference should therefore be investigated rather than presumed.
Separate effect differences from differences in follow-up
Efficacy and effectiveness studies sometimes measure outcomes at different stages. A tightly controlled trial might focus on immediate response, while a pragmatic evaluation examines outcomes over a year.
Before attributing differences to setting or implementation, ask whether the outcomes were measured at comparable times . An apparent efficacy-effectiveness gap may partly be a short-term versus long-term effect.
04 · A Practical Example
How can an intervention work in a trial but perform differently in routine practice?
Hypothetical Example
An adaptive tutoring program moves from a controlled study into routine teaching
Imagine an adaptive tutoring system evaluated in several hypothetical university studies. Early trials use carefully selected courses, instructors receive intensive training, students are reminded repeatedly to use the system, technical support is immediately available, and usage is closely monitored. Later studies evaluate the same system across ordinary courses with routine faculty support and no special adherence program.
Identify the efficacy question The earlier trials estimate whether the tutoring system can improve learning outcomes when implementation and engagement are strongly supported.
Identify the effectiveness question The later studies estimate what happens when institutions adopt the system under conditions closer to normal teaching practice.
Compare implementation conditions The reviewer examines student eligibility, instructor training, required versus voluntary use, technical support, adherence, comparator conditions, and outcome timing rather than simply sorting papers by whether authors called them pragmatic.
Interpret any difference carefully Suppose controlled trials show larger improvements than routine-practice studies. The review can report that effects were larger under more intensively supported conditions while investigating whether adherence, implementation, population, comparator, or other differences plausibly account for the pattern.
Preserve both conclusions The controlled evidence informs whether the system can produce an effect. The pragmatic evidence is more directly informative about the effects institutions might expect under the routine conditions actually studied.
Neither evidence stream needs to invalidate the other. Together they can reveal something more useful: how intervention effects depend on the conditions under which the intervention is delivered.
06 · What This Means for You
Synthesize the conditions of the effect, not only the effect size
When extracting evidence, record features that influence applicability: participant selection, setting, provider expertise, intervention flexibility, adherence support, comparator, outcome relevance, and follow-up. These details help determine what each estimate can reasonably tell you about routine practice.
A simple decision framework
If studies evaluate the intervention under tightly controlled conditions
Interpret their estimates primarily in relation to the conditions and populations actually studied.
If studies deliberately approximate routine practice
Examine how closely their participants, settings, providers, comparators, and implementation resemble the practical decision of interest.
If efficacy and effectiveness estimates differ
Investigate population, adherence, implementation, comparator, outcome, timing, and methodological differences before attributing the gap to any one factor.
If the evidence comes from both types of conditions
Consider presenting them as complementary evidence streams when a single pooled estimate would conceal important differences in applicability.
Your conclusion can then be more precise than “the intervention works” or “the intervention does not work.” You may be able to say that the intervention demonstrates efficacy under supported conditions, while evidence under routine conditions suggests a smaller, similar, larger, or more uncertain effect.
That distinction gives readers information they can actually use when deciding whether evidence is likely to transfer to their own setting.
07 · A Quick Checklist
Before combining efficacy and effectiveness evidence, check:
For each study, verify:
Examine who was eligible and how closely participants resemble the intended real-world population.
Compare study settings with the settings in which the intervention would ordinarily be used.
Record how intensively intervention delivery, provider behavior, and participant adherence were supported or monitored.
Determine whether the comparator represents a realistic alternative in routine practice.
Check whether outcomes and follow-up periods are meaningful for practical decision-making.
Do not classify a study as wholly efficacy-oriented or wholly pragmatic without examining its individual design features.
Investigate reasons for differences between controlled and routine-practice estimates rather than assuming effects must become smaller.
State explicitly which populations and conditions your final conclusion most directly represents.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation