Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Synthesize Evidence When Outcomes Are Measured at Different Times?

The same outcome measured at different times may represent meaningfully different intervention effects. Learn when time points can be grouped, when they should remain separate, and how to avoid selecting follow-up results opportunistically.

560
Synthesizing Outcomes Measured at Different Times Guide 560 of 899
01 · The Question

What if every study measures the same outcome, but not at the same time?

You have several studies evaluating the same intervention and reporting the same outcome. One measures it immediately after treatment. Another reports results at three months. Others provide six-month, one-year, or even longer follow-up.

They may all call the outcome by the same name, but are they estimating the same thing?

Not necessarily. Intervention effects can emerge, strengthen, weaken, disappear, or even reverse over time. Combining measurements taken at meaningfully different stages can therefore produce an average that obscures the very pattern you need to understand.

The challenge is not simply finding a common time point. It is deciding which differences in timing matter for the research question and designing the synthesis accordingly.

02 · The Short Answer

Treat outcome timing as part of the effect being estimated

In Brief

When studies measure outcomes at different times, define clinically or substantively meaningful time frames and synthesize results within those periods only when the measurements represent sufficiently comparable stages of the intervention and outcome process.

Do not automatically pool the nearest available measurements or select whichever follow-up produces the clearest result. Timing decisions should ideally be prespecified, justified by the research question, and handled in ways that avoid counting dependent measurements from the same participants as independent evidence.

03 · What You Need to Know

Time is part of the research question, not merely a feature of data collection

The same outcome can mean something different at different stages

Suppose an educational intervention improves test performance immediately after a course. That tells you something about its short-term effect. It does not necessarily tell you whether the improvement persists six months later.

Likewise, an intervention may produce little immediate change but have effects that emerge gradually. Other interventions may initially work well and then lose their advantage once support ends.

Short-term effect What happens during or soon after the intervention, when immediate effects may be most visible.
Long-term effect What happens after sufficient time has passed for persistence, decay, delayed effects, adaptation, or other changes to become observable.

Neither time point is inherently superior. Which one matters depends on the outcome, intervention, decision being made, and underlying process of change.

Define meaningful time frames rather than demanding identical dates

Exact matching is often unrealistic. One study may assess an outcome at 12 weeks, another at three months, and another at 14 weeks. Treating these automatically as three fundamentally different time points may make synthesis unnecessarily fragmented.

A practical approach is to define meaningful time frames in advance, such as immediate post-intervention, short term, medium term, and long term. Cochrane guidance explicitly allows review authors to group measurements into prespecified intervals when studies are expected to assess outcomes at different times.

The boundaries should come from the substantive problem, not from a universal calendar. “Long term” for postoperative pain might represent a very different interval from “long term” for educational attainment or prevention of a chronic condition.

Timing situation Possible synthesis approach Main question
Measurements differ slightly but represent the same substantive stage Group within a prespecified time frame Are the timing differences unlikely to alter interpretation?
Measurements represent clearly different stages Separate syntheses Could the intervention effect plausibly change over this interval?
Studies report several eligible measurements within one time frame Use a prespecified selection rule or a method accounting for dependency How will one study avoid contributing multiple correlated estimates as independent evidence?
Timing itself is scientifically important Examine effects across time rather than forcing one summary Does the research question concern onset, persistence, decay, or delayed effects?

Choose time frames for substantive reasons

A useful time frame should correspond to a meaningful stage in the outcome or intervention process. Consider how quickly the intervention is expected to work, how long it is delivered, whether effects should persist after it ends, and when the outcome matters to decision-makers.

Cochrane recommends focusing on time points that are meaningful for the decisions the review is intended to inform and, where relevant, considering accepted time points or core outcome sets within the field.

Very narrow windows can leave too few studies for useful synthesis. Very broad windows can make the resulting estimate difficult to interpret. A category such as “long term” is not helpful if it combines measurements at six months and five years when substantial change is expected between them.

Do not choose follow-up times after seeing which results look favorable

Many studies report the same outcome repeatedly. A trial might provide results immediately after treatment, at three months, six months, and one year. If reviewers choose among these after examining the effect sizes or statistical significance, the synthesis becomes vulnerable to selective decision-making.

Outcome timing should therefore be prespecified where possible. Cochrane recommends defining time points or time frames in advance, partly to reduce the possibility that choices are influenced by the observed results.

If your protocol specifies “the measurement closest to six months,” apply that rule consistently. If it specifies the longest available follow-up, use the longest follow-up regardless of whether the result is more or less favorable.

Watch Out

Choosing the follow-up time with the largest effect, smallest P value, or most convenient complete data is not a neutral synthesis decision. When several eligible measurements exist, use a prespecified rule that does not depend on the observed result.

Multiple measurements from the same participants are not independent studies

If one trial reports outcomes at three, six, and twelve months, those three estimates arise partly from the same participants. Treating them as three independent studies in an ordinary meta-analysis gives that study inappropriate influence and ignores the statistical dependence among its estimates.

Cochrane specifically cautions that results from multiple follow-up periods within the same study cannot simply be entered into a standard meta-analysis as independent observations. Options include separate analyses for different time periods, selecting one prespecified time point, or using analytical methods that appropriately account for repeated measurements.

This is one reason time frames are useful. A study might contribute once to a short-term synthesis and once to a long-term synthesis, with those analyses interpreted separately rather than pretending that both estimates are independent within one pooled calculation.

The longest follow-up is not automatically the most informative

Choosing the longest available follow-up from every study sounds objective, and it can be a legitimate prespecified strategy. But it creates another issue: “longest” may mean three months in one study and three years in another.

Those measurements may represent different stages of the intervention's effect. Longer follow-up can also involve greater attrition, changes in exposure, additional treatments, or other events that complicate interpretation.

The appropriate question is not simply, “What is the latest measurement available?” It is, “Which time point best represents the effect relevant to this synthesis?”

Timing can explain apparent disagreement among studies

Suppose short-term studies consistently show a benefit while long-term studies show little difference. It would be misleading to describe the evidence merely as “inconsistent.”

The findings may instead indicate that the intervention produces an initial effect that diminishes over time. That temporal pattern can be more informative than an overall average across all follow-up periods.

Likewise, studies can appear inconsistent because their comparison groups differ. Before interpreting statistical heterogeneity as unexplained disagreement, examine whether studies are actually estimating effects under comparable conditions and at comparable times.

Timing should remain visible in the final interpretation

If your evidence supports an effect immediately after intervention but not at twelve months, saying simply that “the intervention is effective” loses important information.

State the time frame: the intervention improved the outcome in the short term, while evidence of sustained benefit at longer follow-up was limited, uncertain, or absent, depending on what the evidence actually shows.

Cochrane's guidance for Summary of Findings tables similarly recognizes that multiple time points can matter when the result or decision is likely to vary over time.

04 · A Practical Example

How should different follow-up times be handled in practice?

Hypothetical Example

Does a research-writing workshop improve students' writing performance?

Imagine a review of eight hypothetical trials evaluating a research-writing workshop. Studies assess writing performance immediately after the workshop, between two and four months later, between six and nine months later, or at approximately one year. Several studies report more than one of these measurements.

Define the meaningful stages Before examining effect sizes, the reviewer defines immediate post-intervention, short-term follow-up of up to four months, medium-term follow-up from more than four to nine months, and long-term follow-up beyond nine months. These boundaries reflect the review's interest in immediate learning and subsequent retention.
Assign measurements to the time frames A result at three months enters the short-term category. A seven-month assessment enters the medium-term category. A twelve-month assessment contributes to the long-term synthesis.
Handle multiple measurements consistently If a study reports both two- and four-month results within the same short-term category, the reviewer applies a prespecified rule, such as selecting the measurement closest to the end of the defined interval, rather than choosing whichever result is larger.
Synthesize each period separately The reviewer estimates short-, medium-, and long-term effects separately rather than averaging all follow-up measurements into one overall effect.
Interpret the trajectory If benefits are substantial immediately after the workshop but progressively smaller at later follow-up, the review describes evidence of short-term improvement with weaker evidence of sustained effects rather than calling the studies contradictory.

The important result may therefore be the pattern across time, not a single grand average.

05 · What Researchers Often Get Wrong

Timing decisions that can distort an evidence synthesis

Misconception

The same outcome can always be pooled regardless of when it was measured

Outcome identity does not guarantee temporal comparability. An anxiety score measured immediately after treatment and the same score measured two years later can represent meaningfully different intervention effects.

Misconception

You should always use the longest available follow-up

The longest follow-up can be appropriate for some questions, but not universally. It may differ substantially across studies and may not correspond to the time period most relevant to the decision being made.

Misconception

Each follow-up measurement can be treated as another independent result

Repeated measurements from the same participants are statistically dependent. Entering them as independent observations in an ordinary meta-analysis can give a study excessive weight and produce incorrect precision.

Misconception

Time categories such as short term and long term have universal definitions

They do not. Meaningful boundaries depend on the intervention, outcome, field, disease or developmental process, and decision context. Define and justify them for the review rather than importing arbitrary labels.

Misconception

A fading effect means the intervention never worked

An intervention can produce a genuine short-term effect that is not sustained. Persistence is a separate feature of the effect. A synthesis should show that trajectory rather than retroactively erasing the earlier result.

06 · What This Means for You

Decide what time means before deciding what to pool

During protocol development, specify which time periods matter for each important outcome and why. When exact matching is unrealistic, define ranges broad enough to accommodate plausible reporting differences but narrow enough to retain substantive meaning.

A simple decision framework

If measurements occur at slightly different times but represent the same meaningful stage
Group them within a prespecified time frame if the remaining requirements for synthesis are satisfied.
If timing could materially change the meaning or magnitude of the effect
Conduct separate syntheses for the relevant periods.
If one study reports several eligible measurements within the same period
Apply a prespecified selection rule or use a method that accounts appropriately for dependent estimates.
If short- and long-term findings differ
Interpret the temporal pattern rather than reducing the disagreement to a single average.

Record the exact measurement time during data extraction even if you later assign results to broader categories. Cochrane recommends extracting timing information for outcomes explicitly, because it is part of understanding what each result represents.

Finally, keep the time frame visible when you write the conclusion. Readers making decisions usually care not only whether something changes, but also when the change appears and whether it lasts.

07 · A Quick Checklist

Before combining outcomes measured at different times, check:

For each outcome and follow-up period, verify:
Record the exact timing of outcome measurement for every eligible result.
Define substantively meaningful time frames before examining which results are most favorable.
Justify time-frame boundaries using the intervention, outcome process, field conventions, or decision context.
Check whether measurements grouped within one interval genuinely represent a comparable stage of follow-up.
Use a prespecified rule when a study reports several eligible measurements within the same time frame.
Do not treat repeated measurements from the same participants as independent studies in a standard meta-analysis.
Examine whether apparent heterogeneity corresponds to short-, medium-, or long-term differences.
State the relevant follow-up period when reporting the final effect.
08 · Frequently Asked Questions

Common questions about outcome timing in evidence synthesis

How close do follow-up times need to be before I can combine them?

There is no universal numerical threshold. The relevant question is whether the measurements represent a substantively comparable stage of the intervention and outcome process. Define defensible time windows for the particular review rather than applying an arbitrary tolerance.

Should I use the time point reported by the largest number of studies?

Data availability can be considered, but it should not be the only criterion. The selected time point should also be meaningful for the review question. Choosing solely to maximize the number of studies can produce a precise answer to a less useful question.

Can the same study contribute to both short- and long-term meta-analyses?

Yes, when those are separate analyses addressing different time periods. The estimates should not simply be treated as independent observations within the same standard meta-analysis.

What if a study reports an outcome at many time points?

Use the selection rule specified in the review protocol for each time frame, or an analytical approach that properly accounts for repeated correlated measurements. Avoid selecting the time point according to the observed effect size or statistical significance.

Can I combine end-of-treatment measurements when interventions have different durations?

Possibly, but “end of treatment” may occur at very different elapsed times and may represent different stages of change. Consider whether intervention completion itself or time since baseline is the more meaningful basis for comparability.

What if the intervention works initially but the effect disappears later?

Report that temporal pattern. Evidence of a short-term effect and evidence about whether that effect persists answer related but distinct questions. Averaging them can conceal an important finding about durability.

Does measuring outcomes later always provide stronger evidence?

No. Longer follow-up may be essential for durability or delayed outcomes, but it can also involve greater attrition, additional exposures, treatment changes, or other complications. The most informative timing depends on the question being asked.

09 · The Bottom Line

Do not average away the time course of an effect

The Bottom Line

When outcomes are measured at different times, synthesize measurements together only when they represent a substantively comparable time frame; otherwise, preserve short-, medium-, and long-term effects as separate estimates.

Prespecify meaningful time windows and rules for selecting among multiple measurements, account for dependence when the same participants contribute repeated outcomes, and report timing explicitly. Sometimes the most important conclusion is not whether an effect exists on average, but whether it appears, persists, fades, or emerges later.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes