01 · The Question
Can You Search So Broadly That the Evidence Stops Making Sense?
Researchers are often warned against narrow searches, and for good reason. Missing relevant evidence can distort a review. The natural response is to broaden the search: add more synonyms, remove restrictions, include more populations, accept more study designs, search more contexts, and retrieve anything that might conceivably matter.
Eventually, you may have 20,000 records.
The problem is not merely that screening takes longer. If the underlying question and eligibility boundaries are also too broad, the retrieved literature may represent fundamentally different populations, interventions, phenomena, outcomes, settings, and study designs. Combining all of that evidence under one question can make interpretation increasingly difficult.
Broad searching can be methodologically sensible. Undefined breadth is something else.
03 · What You Need to Know
Broad Retrieval and Broad Eligibility Are Not the Same Problem
A search can be broad while the review remains focused
This distinction prevents a great deal of confusion.
Broad retrieval
The search deliberately retrieves many potentially relevant and irrelevant records so that eligible studies are less likely to be missed.
Broad eligibility
The review itself admits many different populations, phenomena, interventions, outcomes, contexts, or designs into the evidence base.
A systematic review may use a deliberately sensitive search that retrieves thousands of irrelevant records, then apply tightly defined eligibility criteria during screening. That can be entirely appropriate.
Cochrane explicitly recommends maximizing search sensitivity while striving for reasonable precision and acknowledges that high sensitivity commonly produces relatively low precision.
The interpretive problem is therefore not simply “too many search results.” It is whether the review has enough conceptual structure to determine what belongs and how different forms of evidence should be understood.
Low precision creates workload even when interpretation remains sound
Precision is the proportion of retrieved records that are relevant. A low-precision search may retrieve large amounts of material that must eventually be excluded.
For example, a search retrieving 10,000 records with 2% precision would contain about 200 relevant records and 9,800 irrelevant ones, assuming those values were known. The problem is obvious: someone must distinguish the 200 from the 9,800.
That workload does not necessarily invalidate the search. Cochrane notes that systematic-review searches often tolerate relatively low precision because sensitivity is important, and title and abstract records can generally be screened much faster than full texts.
Still, poor precision consumes time, increases screening costs, and can make a search operationally difficult. Guidance on peer review of search strategies similarly recognizes that inadequate precision can increase the resources required to complete a review.
More synonyms do not necessarily make a search conceptually broader
Suppose one concept is university students. Adding college students, undergraduates, relevant controlled vocabulary, spelling variants, and terminology used in different jurisdictions may increase retrieval without changing the conceptual population at all.
That is usually desirable expansion within a concept.
By contrast, changing the population from university students to all learners, or expanding an intervention from generative AI to all educational technology, changes the conceptual scope.
Vocabulary breadth
Using multiple legitimate ways of expressing the same concept to improve retrieval.
Conceptual breadth
Expanding what phenomena, populations, interventions, or other substantive categories count as relevant.
Cochrane recommends using a wide variety of terms combined with OR within a concept. That type of breadth is intended to improve sensitivity rather than loosen the research question.
Broad concepts can generate large amounts of noise
Some terms are intrinsically ambiguous. Searching for engagement, performance, technology, learning, AI, or assessment without sufficient conceptual context can retrieve records from very different domains.
The solution is not automatically to delete broad terms. Some may be essential synonyms. Instead, inspect what they retrieve. If one term contributes enormous numbers of irrelevant records without identifying relevant material that other terms miss, it may need qualification or removal.
Search development is iterative. Campbell guidance similarly describes systematic searching as a process in which terms may be modified as retrieval is examined, while recognizing the sensitivity-precision trade-off and diminishing returns from additional searching.
Interpretation becomes harder when the evidence answers different questions
The deeper problem emerges when broad eligibility creates an evidence base containing studies that are technically related but substantively difficult to compare.
Imagine a review asking whether “technology improves education.” Its eligible evidence might include preschool robotics, university learning-management systems, virtual-reality surgical training, mobile language-learning apps, generative AI writing assistants, and workplace simulation.
A very comprehensive search could retrieve all of them. The harder question is what a single synthesis of those studies would mean.
Differences are not inherently a problem. Heterogeneity is normal in research. The issue is whether the differences are compatible with the inferential claim the review intends to make.
This is why defining what you are actually looking for before opening a database matters. The search needs an intellectual boundary, not merely Boolean syntax.
Population breadth can conceal important differences
A review of “students,” for example, might combine primary pupils, adolescents, undergraduates, postgraduate students, and adult professional learners. Sometimes that breadth is appropriate. In other questions, developmental stage, educational context, assessment practices, or learner autonomy could plausibly modify the phenomenon under investigation.
Rather than narrowing automatically, determine whether those groups can legitimately contribute to the same answer. If they cannot, the population boundary may need refinement or the evidence may need to be analyzed in meaningful subgroups.
Study-design breadth can mix different kinds of inference
A review may contain randomized trials, cross-sectional surveys, qualitative interviews, cohort studies, case studies, and mixed-methods research. Such diversity can be entirely appropriate for some review questions.
But these designs do not necessarily estimate the same thing. An experiment estimating an intervention effect and an interview study examining participant experiences contribute different forms of evidence.
The question is therefore not whether diverse designs can coexist, but whether the synthesis respects what each design can support. Deciding which study designs actually matter to the question can prevent comprehensiveness from becoming conceptual indiscrimination.
Broad outcome definitions can produce apparent agreement that means very little
Consider a review in which “student success” includes examination scores, satisfaction, attendance, self-efficacy, course completion, writing quality, and perceived usefulness. All are potentially legitimate outcomes, but they are not interchangeable.
Reporting that most studies found a “positive effect on student success” could conceal substantial differences in what was actually measured.
Broad outcome coverage may therefore require explicit outcome domains and separate synthesis rather than collapsing every favorable result into one category. This is one reason important outcomes should be conceptually defined before the findings are known.
A large result count does not prove that the search is too broad
Some topics simply have enormous literatures. Others are described with ambiguous terminology. Searching multiple databases can produce duplicates. A highly sensitive strategy may intentionally retrieve many irrelevant records.
Therefore, there is no universal result-count threshold at which a search becomes “too broad.”
Watch Out
Do not narrow a search simply because the result count looks intimidating. First determine why the search is large. The problem may be an ambiguous term, unnecessary conceptual breadth, duplicate retrieval across databases, or simply a genuinely large evidence base.
The goal is reasonable precision, not perfect precision
A search returning only eligible studies would be pleasant to screen, but pursuing that ideal too aggressively can sacrifice sensitivity.
Empirical information-retrieval studies illustrate the tension. Montori and colleagues found that MEDLINE strategies optimized for very high precision retrieved substantially fewer systematic reviews than highly sensitive strategies. More recent methodological work likewise describes search-string development as balancing sensitivity against precision rather than maximizing either property independently.
The practical objective is therefore a search broad enough to protect relevant evidence but focused enough that retrieval remains useful and screening remains feasible.
04 · A Practical Example
When a Broad Search Is Useful and When the Scope Is the Real Problem
Hypothetical Example
Searching for AI in education
A researcher searches for studies about artificial intelligence and learning. The strategy includes broad AI terminology, broad education terminology, and numerous synonyms. It retrieves 18,000 records.
First diagnosis The researcher assumes the search is too broad because the result count is large.
Closer inspection Many records concern intelligent tutoring systems, automated assessment, machine learning, generative AI, learning analytics, educational robotics, and administrative AI. These are not merely irrelevant records caused by bad terminology. They represent the breadth of the concept “AI in education.”
Question reconsidered The researcher's actual interest is how university students use generative AI during academic writing. The original question, not merely the syntax, was broader than the intended inquiry.
Search revision The search is rebuilt around the more defensible concepts of higher-education students, generative AI, and academic writing, while retaining multiple synonyms within each concept.
Result Retrieval becomes more focused because the intellectual scope is clearer, not because arbitrary limits were added to force the number down.
This distinction is important. Sometimes you need a better search strategy. Sometimes you need a better question. Database syntax cannot rescue an inquiry whose boundaries remain conceptually undefined.