03 · What You Need to Know
How Expectations Can Shape an Apparent Pattern
Confirmation Bias Is Broader Than Deliberately Ignoring Evidence
Confirmation bias is commonly understood as seeking or interpreting evidence in ways that favor existing beliefs, expectations, or a hypothesis under consideration. Nickerson's influential review describes the phenomenon as appearing in multiple forms rather than as one simple behavior.
That distinction matters for researchers. You do not need to consciously suppress contradictory evidence to be affected by confirmation bias. Expectations can influence which search terms seem natural, which papers attract attention, which methodological weaknesses seem serious, how ambiguous results are classified, and which explanations come readily to mind.
A researcher can therefore be careful, sincere, and still interpret a literature asymmetrically. Bias does not require bad faith.
Expectations Are Not the Problem by Themselves
Researchers routinely begin with theories, prior evidence, hypotheses, and informed expectations. Science would be rather inefficient if every investigation required complete intellectual amnesia.
The problem arises when an expectation becomes difficult to challenge because the procedures for finding or interpreting evidence change depending on whether the evidence supports it.
For example, imagine that you accept a supportive study despite moderate limitations but dismiss a contradictory study because of comparable limitations. The methodological criticism may be legitimate in both cases. The asymmetry is the problem.
Having an expectation
Beginning with a hypothesis, theory, prediction, or interpretation that could be supported or challenged by evidence.
Protecting an expectation
Allowing search, selection, evaluation, or interpretation standards to shift in ways that systematically favor the expected conclusion.
Searching Can Produce the Pattern Before Interpretation Begins
Your evidence base depends partly on how you searched for it.
Suppose you expect educational technology to improve student engagement. A search built mainly around phrases such as “benefits of educational technology,” “technology improves engagement,” and “positive effects of digital learning” is already tilted toward supportive terminology.
A broader search might include neutral descriptions of the intervention and outcome rather than presumed effect direction. It might also search multiple databases, registers, reference lists, and other appropriate sources.
Cochrane emphasizes thorough, objective, and reproducible searching because selective retrieval can produce an unrepresentative evidence base. Its guidance explicitly treats comprehensive searching as one means of minimizing bias in systematic reviews.
For formal systematic reviews, PRISMA 2020 also requires transparent reporting of information sources, full search strategies, eligibility criteria, and study-selection procedures. These practices do not guarantee neutrality, but they make selective decisions more visible and reproducible.
Define Inclusion Rules Before You Know Which Studies Help You
Flexible inclusion criteria create an obvious opportunity for expectations to influence the evidence base.
Imagine discovering a supportive study with a six-week follow-up and deciding that six weeks is sufficient. Later, you find a contradictory study with a five-week follow-up and decide that the follow-up is too short. Each decision can be defended separately. Together, they look rather less innocent.
Prespecifying important eligibility criteria reduces this flexibility. PRISMA recommends reporting the criteria used to determine eligibility, while Cochrane guidance emphasizes defining criteria for including studies and grouping them for synthesis.
Prespecification is especially useful for decisions that could plausibly change once you know the results.
Define What Counts as Support Before Classifying Findings
Researchers sometimes decide that a pattern exists by sorting studies into categories such as supportive, mixed, and contradictory. Those categories can be surprisingly elastic.
Suppose a study finds the predicted association for two outcomes but not three others. Is that supportive? Mixed? Mostly null? The answer depends on what was expected and which outcomes matter.
If possible, establish your interpretive rules before inspecting the full pattern. Specify the relevant outcome, direction, time point, population, comparison, and magnitude. If several outcomes or analyses are legitimate, report that multiplicity rather than selecting whichever one best fits the expected story.
PRISMA 2020 specifically asks systematic-review authors to define the outcomes for which data were sought and explain how results were selected when multiple compatible measures, time points, or analyses exist.
Compare Effect Estimates, Not Just Convenient Labels
Labels such as “significant,” “nonsignificant,” “supports,” and “does not support” can make a literature look more categorical than it is.
Two studies may estimate almost identical effects while falling on different sides of a statistical significance threshold. Conversely, two statistically significant findings may differ substantially in magnitude.
If you expect a positive effect, it is easy to count every significant positive result as confirmation while treating nonsignificant results as merely underpowered. Sometimes that interpretation is justified. Sometimes it is an asymmetrical rescue operation.
Examine effect estimates, uncertainty, methodological quality, and compatibility across studies before assigning interpretive labels.
Use the Same Methodological Standards for Friendly and Unfriendly Results
This is one of the simplest checks and one of the hardest to apply consistently.
When a study supports your expectation, list its weaknesses. When another contradicts it, list its strengths. Then reverse the exercise.
Would you consider the supportive study's sample too small if it had produced the opposite result? Would you describe the contradictory study's measurement tool as well validated if its estimate favored your theory? Would the same amount of missing data concern you equally?
Methodological quality should influence evidential weight, but the standards should not depend on whether the conclusion is convenient. This becomes especially important when deciding whether a minority of stronger studies should receive more evidential weight than a majority of weaker studies.
Actively Search for Evidence That Could Falsify Your Interpretation
A useful discipline is to ask what you would expect to observe if your preferred explanation were wrong.
Then look for that evidence.
If you think an intervention works because of a particular mechanism, identify evidence that could distinguish that mechanism from alternatives. If you think a relationship generalizes, deliberately search for populations in which it might plausibly weaken. If you think a result is robust, inspect studies using methods that do not share the same vulnerability.
This changes the task from collecting confirmations to exposing the explanation to meaningful opportunities to fail.
Look for the Inconvenient Studies Deliberately
Do not assume that the first search results, most highly cited papers, or most familiar reviews represent the complete literature.
Try searches that include neutral terminology and plausible alternative explanations. Examine reference lists and forward citations. Search for null findings, replications, critiques, alternative mechanisms, and relevant grey literature when appropriate to the review question.
Cochrane recommends searching multiple sources and cautions that relying on a single database can retrieve an unrepresentative subset of relevant reports. It also notes that publication status restrictions can contribute to publication bias.
This is particularly important when citation networks make one interpretation appear more dominant than the independent evidence beneath it.
Check Whether Apparently Confirming Studies Are Actually Independent
An expected pattern becomes especially convincing when it seems to recur repeatedly. Before taking comfort in the count, check what is being counted.
Several papers may use overlapping participants, the same cohort, or the same public dataset. Different studies may repeatedly use the same instrument or design. Multiple papers may cite the same original finding without contributing new empirical evidence.
If you already expect the pattern, these distinctions are particularly easy to overlook because every additional supportive paper feels like another confirmation.
Determine whether the studies actually provide independent datasets or repeated analyses of the same evidence.
Ask Whether Agreement Could Come From a Shared Weakness
Suppose six studies all support your hypothesis. Before declaring convergence, identify what all six have in common.
Do they use the same self-report measure? The same cross-sectional design? Similar convenience samples? The same analytical assumptions?
If a shared feature could plausibly produce the finding, agreement may be less independent than it appears. The appropriate question becomes whether the studies could share a weakness capable of generating the same result.
Confirmation bias and correlated methodological bias can reinforce each other rather effectively, which is not the kind of interdisciplinary collaboration anyone requested.
Use Convergence to Challenge Your Preferred Explanation
Evidence from genuinely different methods can help because your explanation must survive different sources of potential error.
If a pattern appears in self-report, behavioral data, longitudinal evidence, and an appropriate experimental design, one measurement-specific or design-specific artifact becomes a less complete explanation.
But methodological diversity should itself be evaluated critically. The methods must address sufficiently related claims, be credible on their own terms, and differ in ways relevant to plausible alternative explanations.
This is the distinction between genuine convergence and repetition of essentially the same evidence.
Independent Reviewers Can Reduce Some Opportunities for Selective Judgment
In systematic reviews, having more than one person independently screen records or collect data can make individual judgment calls more visible. PRISMA 2020 specifically asks authors to report how many reviewers screened records and reports and whether they worked independently.
This does not make reviewers unbiased. Two people can share the same expectation. Independence nevertheless creates a procedure for detecting disagreements that one researcher working alone might never notice.
Where practical, similar logic can be applied outside formal systematic reviews. Ask a colleague who is not invested in the interpretation to examine ambiguous classifications, coding decisions, or apparently anomalous studies.
Document Decisions That Changed After Seeing the Evidence
Research plans sometimes need to change. New terminology appears. An outcome turns out to be measured differently than expected. A planned subgroup proves impossible to analyze.
The problem is not revision. The problem is invisible revision.
Record important changes and why they occurred. Distinguish decisions made before seeing the relevant results from those made afterward. Transparency does not eliminate hindsight, but it prevents hindsight from quietly masquerading as foresight.
Your Preferred Theory Should Have a Losing Condition
Ask yourself a blunt methodological question: what evidence would make me reduce confidence in this interpretation?
If no plausible result would change your view, the interpretation is difficult to test meaningfully.
A useful synthesis identifies not only evidence compatible with the favored explanation but also observations that would count against it. Those conditions need not produce an immediate binary rejection. Scientific conclusions are often revised gradually. Still, a hypothesis that can absorb every possible result is not receiving much of a test.
Watch Out
Do not turn “I considered alternative explanations” into another ritual sentence. Name the strongest plausible alternative, identify what evidence would support it, and check whether that evidence is actually present.