01 · The Question
What if the studies look similar but make different kinds of claims?
You are reviewing studies on the same apparent relationship. Several report that an exposure is associated with an outcome. Others explicitly attempt to estimate what would happen to the outcome if the exposure or intervention changed. Their variables may look nearly identical, and their effect estimates may even point in the same direction.
Can you synthesize them as one body of evidence?
The difficulty is that an associational question and a causal question are not interchangeable. An association asks whether variables differ or move together in observed data. A causal question asks about the effect of changing one factor on another outcome, which requires additional assumptions about the comparison being made and the processes that could produce the observed relationship.
If that distinction disappears during evidence synthesis, a review can quietly turn repeated associations into a stronger causal conclusion than the underlying studies justify.
03 · What You Need to Know
The key distinction is the inferential question, not simply the study-design label
Association and causation ask different questions
An associational question asks whether two variables are statistically related in the observed data. Depending on the design, a study might ask whether students who use a particular learning platform more frequently obtain higher examination scores, whether an exposure is more common among people with a particular outcome, or whether two measurements covary.
A causal question is different. It concerns a counterfactual contrast: what would happen to an outcome under one intervention, exposure, or treatment condition compared with what would happen under another? In practical terms, the question might be whether increasing students' use of the platform would change their examination performance compared with some specified alternative.
Associational question
Are exposure X and outcome Y related in the observed data?
Causal question
What would happen to Y if X were changed, compared with a defined alternative?
The distinction matters because the observed association may reflect a causal effect, but it may also reflect confounding, selection processes, measurement problems, reverse causation, or other sources of bias. The statistical association alone does not tell you which explanation is correct.
Do not classify studies as causal or associational solely from whether they are randomized
Randomization provides a particularly strong basis for causal inference because, when properly implemented, it helps create comparable groups with respect to baseline causes of the outcome. This is one reason randomized trials are generally preferred when estimating intervention effects.
But a simple rule such as “randomized equals causal, observational equals association” is too crude. Non-randomized studies can be explicitly designed to estimate causal effects. Modern causal inference approaches may conceptualize an observational analysis as an attempt to emulate a hypothetical randomized experiment, often called a target trial. Such analyses depend on assumptions that must be examined rather than assumed.
For evidence synthesis, therefore, ask what the investigators were trying to estimate and whether their design and analysis plausibly support that interpretation. The design label is informative, but it should not substitute for examining the actual study.
Statistical adjustment does not automatically establish a causal estimate
A regression model containing age, sex, socioeconomic status, baseline scores, and several other covariates may produce an “adjusted association.” That phrase should not quietly become “effect” in your synthesis.
For an observational estimate to receive a causal interpretation, the underlying causal question and assumptions matter. Among the central concerns is confounding: factors related both to the exposure or intervention and to the outcome may create or distort an observed relationship. Appropriate causal analyses therefore require investigators to identify relevant confounding structures and use suitable design and analytical strategies to address them.
Even sophisticated adjustment cannot guarantee that all relevant confounding has been removed. Some confounders may be unmeasured, poorly measured, incorrectly modeled, or unknowable from the available data.
Watch Out
The number of covariates in a model is not a measure of causal credibility. What matters is whether the design, measurements, assumptions, and analysis appropriately address the sources of bias relevant to the causal contrast being estimated.
Start by identifying the estimand or inferential target
Before combining results, ask what quantity each study is trying to estimate. Two papers may both report a risk ratio, odds ratio, regression coefficient, or mean difference while targeting different scientific questions.
For each study, identify the population, exposure or intervention, comparator, outcome, timing, and intended interpretation. For a causal study, also ask what intervention or exposure contrast is being imagined. “Exposure versus no exposure” may sound straightforward until one study examines initiation, another examines sustained exposure, and another simply compares people according to exposure measured at a single time.
This is also why differences in comparison groups can make apparently similar estimates answer different questions.
Separate the role each kind of evidence plays
Associational evidence can still be highly informative. It may document patterns, identify potentially important relationships, characterize populations, reveal inequalities, generate hypotheses, or establish that a phenomenon deserves further investigation. None of those contributions requires pretending that the evidence estimates an intervention effect.
Causal evidence addresses a different inferential task. It attempts to estimate the consequences attributable to changing an exposure or intervention relative to a specified alternative.
| Question to ask |
Associational evidence |
Causal evidence |
| Primary question |
Are X and Y related? |
What is the effect of changing X on Y relative to an alternative? |
| Interpretive target |
Observed relationship |
Causal contrast |
| Central concern |
Accurate characterization of the observed relationship |
Whether the comparison supports a causal interpretation |
| Role of confounding |
May explain or distort the observed relationship |
Must be addressed sufficiently for the intended causal inference |
| Typical synthesis claim |
Studies consistently report an association |
Evidence estimates an effect under specified assumptions and conditions |
Do not let agreement in direction erase differences in inference
Suppose eight observational studies report positive associations and two randomized trials report positive intervention effects. It may be tempting to say that “ten studies demonstrate a positive causal effect.” That statement changes what eight of those studies established.
A more defensible synthesis preserves the layers of evidence: the observational literature may show a consistent association, while the randomized evidence may provide separate evidence about the effect of an intervention. The two strands can reinforce the coherence of the overall picture without being treated as inferentially identical.
The same principle applies when descriptive and causal evidence answer different questions. Synthesis should preserve those differences rather than making every study contribute to a single undifferentiated conclusion.
Observational causal studies require study-specific risk-of-bias assessment
When non-randomized studies explicitly estimate intervention effects, their causal credibility should be assessed in relation to the particular result being synthesized. Cochrane recommends ROBINS-I for assessing risk of bias in non-randomized studies of interventions. The framework considers domains including confounding, participant selection, intervention classification, deviations from intended interventions, missing data, outcome measurement, and selection of the reported result.
This is more informative than deciding that every observational estimate should receive the same evidentiary weight. Non-randomized studies vary substantially in their capacity to estimate causal effects.
Keep mechanistic evidence separate from evidence of an effect
A plausible mechanism may make a causal interpretation more coherent, but it does not establish that the proposed effect actually occurred in the population under study. Evidence that a mechanism could operate and evidence that an exposure did produce an outcome are different propositions.
When mechanistic studies appear alongside associational and causal studies, preserve that additional distinction rather than using mechanistic plausibility as proof of an observed effect.
Sometimes the appropriate synthesis is parallel rather than pooled
Evidence synthesis does not require every eligible study to contribute to one meta-analytic estimate. If studies target materially different questions, a single pooled effect may have no coherent interpretation even when all studies concern the same broad topic.
You may instead organize the evidence into separate synthesis streams, such as associational studies, randomized causal studies, and non-randomized studies explicitly designed for causal inference. Within each stream, further differences in population, comparator, outcome definition, follow-up period, and risk of bias still need consideration.
This becomes particularly important when outcomes are measured at substantially different times, because even otherwise compatible causal estimates may target different time-specific effects.
04 · A Practical Example
What would this look like in an actual evidence synthesis?
Hypothetical Example
Does frequent use of an online practice system improve examination performance?
Imagine a review containing six hypothetical studies of university students. Three cross-sectional studies report that students who use an online practice system more frequently tend to have higher examination scores. One prospective cohort study measures system use before the examination and reports an adjusted positive association. Two randomized trials assign students to receive either structured access to the practice system or the usual course resources and estimate the effect on subsequent examination scores.
Identify the questions The cross-sectional studies primarily establish an observed relationship between system use and performance. The prospective cohort provides stronger temporal information but does not become causal merely because exposure precedes the outcome. The randomized trials explicitly estimate an intervention effect.
Examine alternative explanations Students who voluntarily practice more may differ in prior achievement, motivation, available study time, digital access, or other characteristics that also predict examination performance. Adjustment may reduce some confounding, but its adequacy must be evaluated rather than assumed.
Synthesize compatible evidence The reviewer describes the observational studies as evidence that greater system use is associated with higher scores. The randomized studies are synthesized separately as evidence concerning whether structured access to the system changes scores relative to the specified control condition.
Interpret the whole evidence base If both strands point in the same direction, the reviewer can report that the associational pattern is consistent with the direction of the experimental evidence. The reviewer should not simply count all six studies as six equivalent demonstrations of a causal effect.
Notice what has been preserved. Every study still contributes useful information, but its contribution is bounded by the question it can credibly address. The synthesis becomes more informative precisely because it refuses to make unlike evidence pretend to be identical.
06 · What This Means for You
Build the synthesis around claims, not merely around study counts
When extracting studies, add an explicit field for the inferential question or intended estimand. Do not wait until the discussion section to notice that half of your studies were answering a different question.
For each result, ask what claim the design and analysis are intended to support. Then decide which results are sufficiently compatible to synthesize quantitatively and which should remain separate but inform the broader interpretation.
A simple decision framework
If a study only characterizes an observed relationship
Synthesize it as associational evidence and keep the language of your conclusion associational.
If a study explicitly attempts causal inference
Identify its causal contrast and assumptions, then assess whether its design, analysis, and risk of bias support that interpretation.
If associational and causal studies address materially different inferential targets
Present separate syntheses or evidence streams rather than forcing them into one pooled estimate.
If different evidence streams point in the same direction
Describe the convergence precisely without implying that every study supplies equivalent causal evidence.
If the evidence streams disagree
Investigate design, confounding, populations, comparators, measurement, timing, and risk of bias rather than resolving the disagreement by counting studies.
Your final conclusion should also distinguish evidence that an intervention changes an outcome from evidence explaining the process through which that change occurs. Evidence about whether something works and why it works can complement each other, but those questions require different inferential claims.
The practical goal is not to create a hierarchy in which one type of study is automatically useful and another automatically worthless. It is to make sure that the strength and wording of each conclusion match the evidence that actually supports it.
07 · A Quick Checklist
Before combining associational and causal evidence, check:
Before synthesizing the studies, check:
Identify the actual question or estimand targeted by each study rather than relying only on its design label.
Determine whether each result supports an associational interpretation, a causal interpretation, or a more limited claim.
For causal claims, identify the intervention or exposure contrast, comparator, population, outcome, and relevant time period.
Examine whether confounding and other important sources of bias have been addressed appropriately.
Do not infer causal credibility merely from statistical adjustment or the number of covariates included.
Check whether apparently similar effect measures actually estimate compatible quantities.
Separate evidence streams when pooling would erase meaningful differences in the questions being answered.
Use language in the final synthesis that preserves the distinction between observed association and estimated causal effect.