01 · The Question
How do you reach a balanced conclusion when an intervention produces both desirable and undesirable outcomes?
An intervention improves the outcome it was designed to change. That sounds encouraging. But participants also experience adverse outcomes, burdens, or complications.
Should the benefits and harms be combined into one overall judgment? Should they be analyzed separately? And what if the evidence about benefits is extensive while harms were measured poorly or reported inconsistently?
These are not minor reporting issues. A synthesis focused only on favorable outcomes can make an intervention appear more attractive than the complete evidence warrants. Yet benefits and harms often cannot simply be subtracted from each other because they differ in importance, frequency, severity, timing, measurement, and certainty.
The task is therefore to preserve both sides of the evidence clearly enough that their trade-offs can be understood.
03 · What You Need to Know
A balanced synthesis requires more than adding a paragraph about side effects
Benefits and harms belong in the same decision, but not necessarily the same calculation
People rarely make intervention decisions based on benefit alone. An intervention may reduce symptoms but cause serious adverse effects. It may improve performance while increasing workload. It may reduce one risk while creating another.
Cochrane recommends that systematic reviews consider adverse as well as beneficial effects, particularly when harms could materially influence treatment or policy decisions.
That does not mean every outcome should be mathematically collapsed into a single benefit-harm statistic. A small improvement in one continuous outcome and a rare serious adverse event are measured on different scales and may carry very different importance for different people.
Benefit-harm synthesis
Brings desirable and undesirable consequences into the same interpretation while preserving what each outcome means.
Single composite score
Combines outcomes mathematically and therefore requires explicit assumptions about how different benefits and harms should be valued and weighted.
The first is generally necessary for balanced interpretation. The second is appropriate only in specific circumstances where its assumptions and value judgments are defensible.
Decide which benefits and harms matter before examining the results
Outcome selection should not be driven by whichever results happen to be available or statistically significant. Important beneficial and adverse outcomes should be specified during protocol development where possible.
Cochrane guidance recommends limiting critical and important outcomes to a manageable set while ensuring that adverse as well as beneficial outcomes are represented. This keeps the synthesis focused on outcomes relevant to decisions rather than on every measurement appearing in the literature.
For harms specifically, reviewers may take a confirmatory approach focused on prespecified adverse effects, an exploratory approach that captures reported adverse effects more broadly, or a hybrid of the two.
| Question |
Benefits |
Harms |
| What should be prespecified? |
Critical intended outcomes |
Known or decision-important adverse outcomes where possible |
| How are outcomes commonly collected? |
Often predefined and measured systematically |
May be predefined, actively monitored, passively reported, or discovered unexpectedly |
| What evidence may be needed? |
Often trials designed around intended outcomes |
Trials plus potentially broader evidence when rare, delayed, or incompletely captured harms matter |
| What does non-reporting mean? |
Outcome may not have been measured or reported |
Cannot safely be interpreted as absence of harm |
Harms may require a broader evidence search
A trial designed to establish benefit may be too small or too short to detect uncommon or delayed harms. Some adverse effects become apparent only after an intervention is used by much larger populations or for longer periods.
Cochrane therefore notes that reviews of adverse effects may require a broader range of evidence sources than reviews focused on beneficial outcomes. Depending on the question, relevant evidence may include non-randomized studies or other sources capable of contributing information about rare or longer-term adverse outcomes.
This does not mean treating every source as equally credible. Different study designs bring different strengths and risks of bias. Rather, the evidence base needed to answer a harms question may legitimately differ from the one used to estimate an intended benefit.
Distinguish an adverse event from an adverse effect
Terminology matters. Cochrane distinguishes an adverse event, meaning an unfavorable or harmful outcome occurring during or after an intervention without necessarily being caused by it, from an adverse effect or harm, where a causal relationship with the intervention is at least reasonably possible.
This matters during synthesis because a table of events occurring in an intervention group does not automatically establish that the intervention caused them. Comparative evidence, timing, biological or behavioral plausibility, study design, ascertainment, and other information may be relevant to causal interpretation.
How harms were measured can matter as much as how many were reported
Adverse outcomes may be actively sought using structured questionnaires, laboratory tests, scheduled assessments, or predefined surveillance. Alternatively, investigators may simply record events participants happen to mention.
Those approaches can produce very different event counts.
Cochrane warns that harms data are often collected and reported less rigorously than primary beneficial outcomes. Before synthesizing adverse effects across studies, examine whether definitions and ascertainment methods are sufficiently comparable.
Watch Out
“No harms were reported” can mean no events occurred, no relevant events were detected, no systematic monitoring was performed, or the events were not reported in the publication. Those interpretations are not equivalent.
Do not silently convert missing harms data into zero events
If a study never mentions a particular adverse outcome, entering a zero into the review dataset can create false reassurance.
Cochrane generally recommends against extracting zero adverse events unless the study explicitly reports zero, and even an explicit zero should be interpreted in light of how rigorously the outcome was monitored.
A study using active surveillance for a predefined harm provides different evidence from a study relying on spontaneous participant reports. The number “0” looks identical in a spreadsheet. Its evidentiary meaning may not be.
Rare harms create special problems
A serious adverse effect can matter even when it occurs infrequently. Yet rare events are precisely the outcomes that individual trials may be poorly equipped to detect.
If no serious events occur in a small trial, the correct conclusion is not necessarily that the intervention cannot cause that event. The study may simply provide little information about a very low event probability.
Precision therefore matters. Confidence intervals and the amount of accumulated exposure can be more informative than whether an event count happened to be zero.
Benefits and harms may occur on different timelines
Benefits can appear quickly while harms accumulate slowly, or the reverse. An intervention might produce immediate symptom relief but create a risk that becomes relevant only after prolonged exposure.
This means that differences in outcome timing should be examined separately for beneficial and adverse outcomes. A three-week trial may provide useful evidence about short-term benefit while telling you very little about a harm that typically develops after a year.
Do not let one overall average conceal conflicting outcomes
An intervention may help the primary outcome while worsening another outcome that matters to participants. Those effects should remain visible.
If the intervention improves one outcome but harms another, the synthesis moves beyond simple “works versus does not work” language. The relevant question becomes how the outcomes compare in importance, magnitude, frequency, certainty, and reversibility.
That situation deserves explicit treatment rather than allowing the favorable primary outcome to dominate the conclusion. The problem becomes especially important when an intervention helps one outcome but harms another.
Do not decide the trade-off using statistical significance
Suppose the benefit estimate is statistically significant while the harm estimate is not. That does not automatically establish that benefits outweigh harms.
A harm estimate may be imprecise because events are rare or studies are small. Conversely, a precisely estimated but trivial benefit can reach conventional statistical significance without being important to decision-makers.
Interpret effect sizes, absolute risks where meaningful, uncertainty, severity, timing, and the importance of the outcomes. Statistical significance does not supply the value judgment required to decide how much harm is acceptable for a particular benefit.
The final balance can depend on values and preferences
Even when benefits and harms are estimated accurately, there may be no universal answer about whether the trade-off is worthwhile.
People can value outcomes differently. A modest benefit accompanied by frequent mild inconvenience may be acceptable to some and unattractive to others. A small probability of a severe irreversible harm may dominate the decision for one person while another facing a serious baseline condition may accept that risk.
Cochrane guidance on interpreting review findings therefore separates description of the evidence from healthcare recommendations and recognizes that decisions can depend on values, preferences, costs, and other factors in addition to the balance of benefits and harms.
04 · A Practical Example
What does a balanced benefit-harm synthesis look like?
Hypothetical Example
A digital study-support program improves completion but increases workload
Imagine a review of a hypothetical digital support program for university students. Across several trials, the program increases assignment completion. The studies also report additional weekly workload, frustration with notifications, withdrawal from the program, and several other undesirable outcomes with varying levels of detail.
Synthesize the intended benefit The reviewer estimates the program's effect on assignment completion using studies with compatible definitions and comparison conditions.
Define the important harms and burdens Additional workload, withdrawal because of the intervention, and substantial distress are treated as separate outcomes rather than collapsed into a generic “negative effects” category.
Examine ascertainment Some trials asked every participant about burden using a standardized questionnaire. Others mentioned complaints only when students volunteered them. The reviewer does not treat those methods as equally capable of detecting adverse experiences.
Synthesize each important outcome Completion, workload, withdrawal, and distress are summarized separately with their own effect estimates and uncertainty where the evidence permits.
Present the trade-off The conclusion states that the program improves assignment completion but also increases workload, while evidence concerning less common adverse experiences remains uncertain because monitoring and reporting were inconsistent.
The reviewer does not declare the intervention simply “beneficial” because the primary outcome improved. Nor does the existence of undesirable outcomes automatically make the intervention harmful overall. The evidence is presented so that the trade-off remains visible.
06 · What This Means for You
Make both sides of the intervention visible before judging the trade-off
Plan benefit and harm outcomes together during protocol development. For each important adverse outcome, determine how it is defined, how it was ascertained, how long participants were observed, and whether the available study designs are capable of detecting it.
A simple decision framework
If important benefits and harms are measured reliably in the same studies
Present both evidence streams side by side while preserving their separate effect estimates and uncertainties.
If trials provide strong benefit evidence but weak harms surveillance
State that asymmetry explicitly and consider whether additional evidence sources are needed for the harms question.
If adverse outcomes use inconsistent definitions or ascertainment methods
Assess whether quantitative pooling remains meaningful rather than combining unlike events merely because they share a broad label.
If no adverse events are reported
Determine whether zero events were explicitly documented and whether surveillance was adequate before drawing any safety conclusion.
If important outcomes favor different choices
Present the trade-off rather than converting multidimensional evidence into a simple declaration that the intervention works or does not work.
Where possible, absolute effects can make the trade-off easier to interpret. Saying that an intervention prevents a certain number of undesirable outcomes per 1,000 people while producing a certain number of additional adverse outcomes per 1,000 can be more informative than presenting relative effects alone, provided the baseline risks used are appropriate.
Most importantly, preserve uncertainty symmetrically. Strong evidence of benefit alongside sparse evidence about harms should not be rewritten as “effective and safe.” Sometimes the accurate conclusion is less satisfying but more useful: benefit is reasonably well established, while important uncertainty about harms remains.
07 · A Quick Checklist
Before drawing conclusions about benefits and harms, check:
For a balanced synthesis, verify:
Include important adverse outcomes as well as intended beneficial outcomes.
Prespecify important known harms where possible while allowing an appropriate strategy for detecting unexpected harms.
Check whether harms were actively monitored, passively reported, or not systematically collected.
Compare definitions and ascertainment methods before pooling adverse outcomes across studies.
Do not code an unreported adverse outcome as zero unless the study clearly supports that interpretation.
Consider whether rare or delayed harms require additional study designs or evidence sources.
Compare benefits and harms using effect magnitude, absolute risk, uncertainty, severity, timing, and importance rather than statistical significance alone.
State explicitly when evidence about harms is substantially weaker or less complete than evidence about benefits.