03 · What You Need to Know
What Search Sensitivity Actually Means
Sensitivity is the proportion of relevant records retrieved
In information retrieval, sensitivity is also commonly called recall. Cochrane defines sensitivity as the number of relevant reports identified by the search divided by the total number of relevant reports in the resource being searched.
The formula is straightforward. Applying it to a real research topic is harder because you usually do not know that the database contains exactly 100 relevant records.
This creates the central problem in evaluating sensitivity: the denominator is usually partly unknown.
A large result set does not prove high sensitivity
A search returning 20,000 records can still miss important evidence. It may contain broad ambiguous terminology that retrieves large quantities of irrelevant material while omitting a synonym, subject heading, spelling variant, historical term, or conceptual representation used by relevant studies.
Conversely, a search returning 1,500 records might retrieve nearly all of the relevant literature for a tightly defined topic.
Volume and sensitivity are therefore different properties.
High retrieval volume
The search returns many records, relevant and irrelevant combined.
High sensitivity
The search retrieves a high proportion of the relevant records that should be found.
This is why simply having thousands of search results does not demonstrate comprehensiveness.
Known relevant studies provide an important diagnostic test
Before finalizing a search, identify studies already known to be relevant. These may come from preliminary searching, previous reviews, expert recommendations, citation searching, reference lists, or other sources.
Then ask whether your strategy retrieves them.
Cochrane recommends checking whether a strategy finds key publications or studies included in similar reviews as a basic test of search performance. If a clearly relevant publication is indexed in the database you are searching but your strategy does not retrieve it, investigate why.
Perhaps the study uses terminology absent from your strategy. Perhaps its indexing suggests another subject heading. Perhaps a filter removes it. Perhaps one of your concept blocks is unnecessarily restrictive.
Each missed study can therefore teach you something about the search.
Retrieving all known studies still does not prove adequate sensitivity
This qualification matters.
A strategy can retrieve every study you already know while missing relevant studies that use different terminology. Cochrane specifically cautions that finding only known records is not sufficient evidence of adequate search performance because the strategy may have become biased toward those familiar studies.
This can happen when search development becomes circular. You inspect five known papers, extract their terminology, modify the strategy until all five appear, and conclude that the search is sensitive. Yet all five may come from the same research tradition and use similar language.
A stronger benchmark set contains varied terminology, publication periods, journals, indexing patterns, and methodological characteristics where possible.
Relative recall can estimate sensitivity when the full universe is unknown
One practical approach is relative recall. Instead of pretending that every relevant study is known, you define a benchmark set of relevant publications identified through independent or complementary methods and determine what proportion your search retrieves.
A 2025 practical guide describes relative recall as an accessible method for estimating search-string sensitivity using a predefined set of benchmark publications.
But the word relative matters. Retrieving 95% of a benchmark set does not prove that you retrieved 95% of every relevant publication that exists. The result depends on how complete and representative the benchmark set is.
Missed benchmark studies are often more informative than the percentage itself
Suppose your search retrieves 47 of 50 benchmark studies. The estimated relative recall is 94%.
Should you immediately add terms until it reaches 100%?
Not necessarily. First inspect the three missed studies.
One might not be indexed in the database. Another might be a conference abstract outside the intended source coverage. The third might reveal a genuinely missing synonym that would retrieve additional relevant literature.
Only the third example necessarily identifies a defect in the search string itself.
Sensitivity assessment should therefore combine quantitative measures with error analysis.
Studies discovered through other methods can test the database strategy
Relevant studies may emerge later through citation searching, reference-list checking, trial registries, supplementary database searches, handsearching, author contact, or other methods.
When that happens, ask a useful retrospective question:
If this record was present in the database at the time of searching, should my database strategy have found it?
If yes, test the record against the strategy. Determine which component failed.
Cochrane notes that citation searching can help assess completeness. If citation searching identifies many additional relevant records that were available in the databases searched but were not retrieved by the original strategy, that may indicate that the original searches should be revisited.
Check whether restrictions are causing sensitivity loss
A search can have excellent subject terminology and still lose relevant evidence because of filters.
Age limits, language restrictions, study-design filters, population filters, date limits, field restrictions, and NOT statements can each remove records.
When sensitivity seems inadequate, test the underlying subject search separately from these restrictions. The problem may not require another synonym. It may require removing a condition that should never have been mandatory.
This is especially important when testing whether a search filter is too restrictive.
Search sensitivity can fail at the concept level
Suppose your strategy has three concept blocks:
Population AND Intervention AND Outcome
A missed study may satisfy the population and intervention blocks but fail the outcome block because its abstract uses unexpected terminology.
Adding another outcome synonym may solve that particular case. Alternatively, repeated failures may reveal that the outcome itself is an unreliable retrieval concept.
In that situation, the better solution might be to leave the outcome out of the search strategy and apply it during screening.
Diagnosing the level at which records disappear is more informative than adding terms indiscriminately.
Search sensitivity also depends on the sources searched
A highly sensitive strategy within one database does not guarantee comprehensive identification of evidence overall.
If relevant studies are not indexed in that database, no amount of keyword refinement can retrieve them there. Comprehensive evidence identification may require multiple databases and supplementary methods depending on the research question.
This means two different questions should be kept separate:
Strategy sensitivity
How effectively does the search expression retrieve relevant records available within a particular information source?
Overall search-process coverage
How effectively does the complete set of databases and supplementary methods identify the relevant evidence universe?
A strategy can perform well within MEDLINE while the overall review still misses studies available only elsewhere.
There is no universal sensitivity threshold
Researchers sometimes want a number: Is 90% enough? 95%? 99%?
There is no universal threshold applicable to every search. The acceptable level depends on the purpose of the search, consequences of missing evidence, feasibility of additional screening, type of evidence, and review methodology.
Cochrane recommends that intervention-review searches aim to maximize sensitivity while striving for reasonable precision. That priority makes sense for reviews intended to identify evidence comprehensively. A rapid review or exploratory literature search may explicitly accept different trade-offs.
| Search purpose |
Typical sensitivity priority |
Practical implication |
| Comprehensive systematic review |
Very high |
More irrelevant retrieval may be accepted to reduce missed evidence |
| Rapid evidence synthesis |
High but potentially constrained |
Explicit methodological shortcuts may trade some coverage for feasibility |
| Scoping or mapping search |
Depends on the objective |
Broad conceptual coverage may matter more than highly focused precision |
| Exploratory background search |
Often lower |
Finding enough representative literature may be sufficient |
Sensitivity should be assessed iteratively
Search evaluation is not necessarily a final inspection after the strategy is complete. Test during development.
When a benchmark study is missed, investigate it. When a new synonym is added, examine what it contributes. When a filter is introduced, check whether relevant records disappear. When citation searching finds additional eligible studies, determine whether the database strategy should have retrieved them.
PRESS provides a complementary form of quality assurance by systematically reviewing translation of the research question, Boolean and proximity operators, subject headings, text words, syntax, and limits and filters. Peer review cannot prove complete sensitivity, but it can identify errors and weaknesses that threaten retrieval.
Watch Out
Do not optimize a search solely until it retrieves every benchmark paper. A strategy overfitted to known studies can appear highly sensitive while remaining poor at finding relevant literature that uses unfamiliar terminology.