03 · What You Need to Know
Testing a Filter Means Measuring What It Costs as Well as What It Saves
Start with the right question
Researchers often evaluate a filter by asking:
How many results did it remove?
That tells you how restrictive the filter is numerically. It does not tell you whether it is appropriately restrictive.
A better question is:
Which relevant records did it remove?
If a filter eliminates 80% of your retrieval while retaining nearly all eligible evidence, the reduction may be highly useful. If it eliminates 80% while losing one quarter of the eligible evidence, the same reduction tells a very different methodological story.
Run the search with and without the filter
The simplest useful test is an A/B comparison.
Run your subject search without the filter and record the result count. Then apply the filter and record the new count.
Suppose:
Unfiltered search = 6,000 records
Filtered search = 1,800 records
The filter removed 4,200 records, or 70% of the original set.
That is useful descriptive information, but the analysis is not finished. The next task is to determine what those 4,200 records contain.
Identify the records that the filter removed
When the database supports set combinations or search histories, isolate the difference between the two result sets.
Conceptually, you want:
Unfiltered results NOT Filtered results
This difference set contains the records excluded by the filter.
You do not necessarily need to screen every excluded record during preliminary search development. Inspecting a purposeful or random sample can reveal why records are being removed and whether potentially relevant studies are present.
If you repeatedly find eligible or highly plausible records in that difference set, the filter deserves scrutiny.
Use known relevant studies as benchmark records
Before finalizing a systematic search, you may already know several relevant studies from scoping searches, previous reviews, citation chasing, expert knowledge, or other discovery methods.
These records can function as benchmark studies.
Run the unfiltered strategy and determine which benchmark studies it retrieves. Then apply the filter and check them again. If known relevant records disappear specifically because of the filter, you have identified a concrete sensitivity problem.
Watch Out
Passing the benchmark test does not prove that the filter is fully sensitive. Known studies may share terminology, indexing, publication years, or other characteristics that make them easier to retrieve than unknown eligible studies. Benchmark testing is diagnostic evidence, not proof of completeness.
A reference set should be defined independently where possible
For a more formal test, assemble a set of studies already judged relevant through methods that are not dependent solely on the filter being evaluated.
This might come from a previous high-quality review, handsearching, multiple complementary searches, citation searching, or a completed screening process.
The filter can then be evaluated according to how many members of that reference set it retrieves.
Search-filter development studies commonly use reference standards and report performance measures such as sensitivity, specificity, and precision. A review of methodological filter evaluations found sensitivity or recall and precision among the measures reported most frequently.
Sensitivity asks how many relevant records survived
For a known reference set, sensitivity can be expressed as:
A sensitivity of 92% in this example means that the filter retrieved 46 of the 50 known relevant records. It also means four known relevant records were missed.
Whether that performance is acceptable depends on the purpose of the search. There is no universal percentage at which every filter becomes "safe." A systematic review aiming for comprehensive retrieval may evaluate a 92% filter very differently from a rapid exploratory search.
Relative recall can provide a practical evaluation
When the complete universe of relevant studies is unknown, researchers can evaluate a search against a predefined benchmark set. This is often described using relative recall.
A recent practical guide defines relative recall as the retrieval overlap between the evaluated search string and a search designed to capture the benchmark publications. It proposes this approach as a comparatively accessible way to estimate search sensitivity for systematic reviews.
The logic is straightforward: if you have a defensible set of relevant publications identified independently, determine what proportion your filtered strategy retrieves.
The result remains relative to that reference set. It should not be interpreted as proof that the strategy retrieves the same proportion of all relevant studies that exist.
Precision tells you what the filter saves
Sensitivity measures retention of relevant records. Precision addresses a different question: among the records retrieved, how many are relevant?
A filter can improve precision substantially while reducing sensitivity. That is not contradictory. It simply means the filter made the result set cleaner while also losing some relevant evidence.
The methodological decision concerns whether that trade-off is acceptable. This is why sensitivity and precision should be considered together.
Do not assume that the most sensitive filter is automatically best
A maximally sensitive filter may retrieve nearly every relevant record while doing little to reduce screening burden. Another filter may sacrifice a small amount of sensitivity in exchange for a major gain in precision.
Research comparing randomized-trial search approaches has demonstrated that filters can substantially reduce retrieval with relatively small losses of included studies in some contexts.
Conversely, other evidence types show much less favorable trade-offs. In diagnostic test accuracy searches, evaluated methodological filters have sometimes missed substantial proportions of relevant studies. One analysis found that sensitive filters still failed to identify between 2% and 28% of relevant studies, while more specific filters missed 39% to 42%.
Filter performance is therefore an empirical property, not something you can infer from the elegance of the search string.
Check why benchmark studies were lost
When a relevant record disappears, diagnose the failure rather than immediately deleting the filter.
The record might lack a required indexing term. Its study design may be described using unexpected terminology. Its participants may fall outside a database age category despite satisfying your protocol. The record may not yet be fully indexed. Or your filter may contain an avoidable Boolean or syntax problem.
Different causes imply different solutions.
| Why a relevant record was lost |
Possible response |
| Missing indexing |
Add appropriate text-word coverage or reconsider an indexing-dependent restriction |
| Unexpected terminology |
Expand or revise the relevant search concept |
| Database category does not match eligibility criterion |
Remove the filter or apply the criterion during screening |
| Filter has known sensitivity limitations |
Consider a more sensitive validated filter or no filter |
| Boolean or syntax error |
Correct the strategy and rerun the comparison |
Test one consequential restriction at a time
If you simultaneously add Humans, English, Adult, randomized trial, and publication-date restrictions and 70% of the records disappear, you will not know which restriction caused the loss of a relevant study.
During search development, introduce consequential filters separately where practical. Record the result count after each change and retest your benchmark records.
This resembles controlled experimentation on the search strategy. Academic methods occasionally reward changing only one thing at a time.
Built-in filters deserve the same testing
A checkbox should not escape evaluation merely because you did not write its syntax yourself.
PubMed explicitly notes, for example, that its age and sex filters rely on MeSH indexing and can exclude relevant records that lack the corresponding MeSH terms.
If you are considering built-in Humans, Age, or Study Type filters, test them in the same way you would test a filter copied into the search box.
Use an appropriate benchmark set
A weak benchmark set can give false reassurance. If all your known relevant studies came from the same journal, use similar terminology, or were discovered by an earlier version of the same search, they may not challenge the filter meaningfully.
Where possible, include diverse relevant records: older and newer publications, different terminology, different journals, different indexing patterns, and examples identified through methods independent of the search under evaluation.
The more representative the reference set is of the eligible literature, the more informative the test becomes. It still remains a sample of what is known, not a guarantee about what remains undiscovered.
Decide in advance what level of loss is acceptable
It is easy to rationalize a filter after seeing how convenient the smaller result set looks. A more defensible approach is to consider your retrieval priorities before evaluating the filter.
For a comprehensive systematic review, even a small loss may deserve investigation. For a rapid review conducted under explicit time constraints, some reduction in sensitivity may be an accepted methodological trade-off. For an exploratory search intended only to identify representative literature, precision may matter more.
There is no context-free threshold that settles this decision. The standard should follow the purpose of the search and should be reported honestly.