03 · What You Need to Know
What Search Precision Tells You and What It Does Not
Precision is the proportion of retrieved records that are relevant
Cochrane defines precision as the number of relevant reports identified divided by the total number of reports identified.
Unlike true sensitivity, precision can be calculated directly once all retrieved records have been assessed for relevance. During search development, however, you may only be able to estimate it from a sample.
Low precision does not automatically mean a bad search
A systematic-review search may deliberately retrieve large numbers of irrelevant records because the priority is avoiding missed studies.
Cochrane explicitly notes that searches for systematic reviews should seek to maximize sensitivity while striving for reasonable precision. Increasing comprehensiveness commonly reduces precision and retrieves more non-relevant reports.
A search with 3% precision may therefore be entirely defensible if it captures evidence that a 20% precision strategy misses.
Context is essential.
Low precision
A relatively small proportion of retrieved records are relevant.
Poor search
A strategy that fails its retrieval purpose because of inappropriate concepts, terminology, logic, restrictions, coverage, or performance.
The two can overlap, but they are not synonyms.
Result count is not precision
A search returning 500 records is not inherently more precise than one returning 5,000.
Suppose Search A retrieves 500 records, of which 10 are relevant. Its precision is 2%.
Search B retrieves 5,000 records, of which 500 are relevant. Its precision is 10%.
Search B is ten times larger but five times more precise.
Result count tells you workload. Precision tells you the concentration of relevance within that workload.
You can estimate precision from a sample during search development
Screening the complete result set merely to decide whether the search is precise enough would defeat much of the purpose of preliminary evaluation.
Instead, inspect a sample of retrieved records and classify them according to a clearly defined relevance criterion.
For example, suppose you randomly inspect 200 records and judge 24 potentially relevant:
A sample-based estimate is less certain than evaluating the complete retrieval, especially if the sample is small or non-random. Still, it can be useful for comparing search versions and diagnosing obvious noise.
Before narrowing the search, identify where the noise comes from
Low precision is a symptom. The useful question is what causes it.
Common causes include ambiguous terms, overly broad truncation, broad subject headings, unnecessary synonyms, concepts searched in very broad fields, weak phrase structure, and terminology with unrelated meanings in other disciplines.
Suppose you search AI as a free-text abbreviation for artificial intelligence. You may retrieve records in which AI means something entirely different. The problem is not that the overall topic is too broad. One term is generating noise.
That suggests a targeted correction rather than adding another mandatory concept to every record.
Look at irrelevant records in clusters
Instead of dismissing irrelevant records one at a time, ask whether they share a pattern.
Perhaps most false positives come from veterinary studies. Perhaps one acronym has another disciplinary meaning. Perhaps a truncation retrieves an unintended word family. Perhaps a broad subject heading is responsible for most of the noise.
Patterns make search refinement more surgical.
| Source of low precision |
Possible refinement |
Risk to check |
| Ambiguous synonym |
Remove it, qualify it, or restrict its field |
Relevant studies that rely on that term may disappear |
| Overly broad truncation |
Use a longer stem or alternative terms |
Useful word variants may be lost |
| Broad concept searched everywhere |
Consider more appropriate title, abstract, keyword, or subject-heading fields |
Relevant records may express the concept only outside the restricted field |
| Unnecessary OR terms |
Test their contribution and remove terms producing noise without useful retrieval |
Rare relevant terminology may be lost |
| Broad subject heading |
Consider a more specific heading or appropriate qualification |
Narrower indexing may reduce sensitivity |
Adding another AND concept is not always the best way to improve precision
If a search retrieves too much irrelevant material, the intuitive response is often to add another concept with AND.
That certainly narrows retrieval, but it creates another condition every relevant record must satisfy. If the new concept is inconsistently described, precision may improve at the cost of substantial sensitivity.
For example, adding an outcome block can reduce 8,000 records to 1,000 while excluding studies that measure the outcome but do not mention it in searchable metadata.
Before adding a new mandatory concept, first consider whether existing concepts can be made more discriminating. This is one reason additional search terms can sometimes make the strategy worse.
Precision should be considered alongside sensitivity
Suppose two strategies perform as follows against a benchmark:
| Strategy |
Benchmark sensitivity |
Precision |
| Search A |
98% |
4% |
| Search B |
82% |
18% |
Search B is much cleaner. But it misses substantially more relevant evidence.
Which is preferable depends on the search purpose. For a comprehensive systematic review, Search A may be more defensible despite its screening burden. For another type of search with different objectives and explicit resource constraints, the balance may differ.
The central issue is explored directly when comparing search sensitivity and precision.
Screening burden is a legitimate consideration
Precision is not merely an aesthetic preference for tidy search results. Low precision consumes researcher time and resources.
Cochrane notes, however, that abstracts can often be screened relatively quickly and estimates a rate of approximately 60 to 120 abstracts per hour, or roughly 500 to 1,000 over an eight-hour period. The point is not that screening thousands of records is trivial. Rather, a seemingly large result set should be considered in relation to the total work of the review and the cost of missed evidence.
A search should not be made methodologically weaker merely to save a modest amount of screening.
Extremely low precision can still signal a design problem
Although low precision can be justified, it should not become an excuse for an indiscriminate search.
If almost every record is irrelevant, inspect why. Perhaps one concept is represented too broadly. Perhaps a generic term is searched without context. Perhaps your Boolean grouping is incorrect. Perhaps the query has been translated badly into another database.
PRESS identifies research-question translation, Boolean and proximity operators, subject headings, text words, spelling and syntax, and limits and filters as areas requiring peer review. Problems in any of these areas can generate avoidable irrelevant retrieval.
Precision may differ considerably between databases
The same conceptual strategy can behave differently across databases because their coverage, controlled vocabulary, indexing, syntax, and user interfaces differ.
A term that is highly discriminating in one database may generate substantial noise in another. A subject heading available in MEDLINE may have no exact counterpart elsewhere.
Therefore, do not assume that a precision improvement achieved in one database can simply be copied into another. This is part of the broader problem of translating a search strategy between databases.
There is no universal precision threshold
Unlike a laboratory reference range, precision does not come with a universal cutoff separating acceptable from unacceptable searches.
A 10% precision rate means that roughly one in ten retrieved records is relevant. Whether that is excellent, mediocre, or impractical depends on the topic, review method, sensitivity achieved, number of records, screening resources, and consequences of missing evidence.
For a narrow clinical question, 10% may appear low. For a difficult interdisciplinary topic with unstable terminology, it may be quite efficient.
The more useful question is whether additional precision can be obtained without losing relevant evidence that matters.
Stop refining when further gains are not worth the sensitivity cost
Suppose a revision raises estimated precision from 6% to 9% while all benchmark studies remain. That may be worthwhile.
Another revision raises it from 9% to 15% but removes four benchmark studies. For a comprehensive review, the second improvement may be difficult to justify.
Search refinement therefore reaches a point of diminishing returns. You are not trying to create a result set in which every record is relevant. You are trying to create a defensible retrieval set in which irrelevant records are manageable and relevant evidence remains discoverable.
Watch Out
Do not optimize precision until the search results look pleasing. Every improvement in cleanliness should be checked for what disappeared along with the noise.