01 · The Question
Why Can the Same Literature Search Produce Very Different Result Counts?
You translate a search into several databases and expect broadly comparable results. One database retrieves 2,000 records, another 7,000, and a third only 900. The discrepancy can look suspicious enough to suggest that one of the searches must be wrong.
Sometimes it is. A mistranslated field code, missing subject heading, unsupported proximity operator, or misplaced Boolean expression can substantially alter retrieval.
But large differences can also be legitimate. Databases do not contain identical collections of records, use identical controlled vocabularies, index documents in the same way, or necessarily expose the same searchable fields. The useful question is therefore not whether the counts match. It is whether each search accurately represents the intended concepts within that particular database.
03 · What You Need to Know
Equivalent Search Logic Does Not Mean Identical Retrieval
Databases contain different literature
The most fundamental reason for different result counts is coverage. Bibliographic databases do not index exactly the same journals, conference proceedings, document types, geographic literature, or publication periods.
Cochrane notes that MEDLINE and Embase overlap, but the degree of overlap varies substantially by topic. It also identifies subject-specific and regional databases as potentially important because they may contain literature not indexed elsewhere.
Consequently, even perfectly translated searches can retrieve different numbers of records because they are searching different collections.
Database difference
The strategies are appropriately translated, but retrieval differs because the databases contain or index literature differently.
Translation problem
The intended search logic has changed during adaptation because syntax, fields, vocabulary, or operators were translated incorrectly.
Controlled vocabularies are not interchangeable
Databases may use different indexing vocabularies. MEDLINE uses Medical Subject Headings (MeSH), while Embase uses Emtree. Cochrane explicitly notes that the controlled-vocabulary terms and indexing approaches for MEDLINE and Embase are not identical.
A MeSH descriptor therefore cannot simply be pasted into an Embase strategy and assumed to have the same meaning or retrieval behavior. The corresponding Emtree concept may occupy a different position in its hierarchy, have different narrower terms, or be indexed differently.
For each database, identify the appropriate controlled-vocabulary term independently and inspect its scope and hierarchy.
Explosion can produce different conceptual breadth
Many controlled vocabularies are hierarchical. Exploding a subject heading generally retrieves records indexed with that heading together with relevant narrower terms beneath it, according to the particular system.
Two apparently corresponding headings can therefore retrieve quite different sets if their hierarchical structures differ.
Cochrane recommends using appropriate controlled vocabulary, including exploded terms where relevant, but this requires database-specific judgment. Matching the visible heading names is not enough.
The same free-text expression may search different fields
A search for a word can mean different things depending on the platform. One interface may search title and abstract by default. Another may include author keywords, subject fields, or additional metadata. A third may apply automatic query processing.
This can produce major count differences even when the text typed into the search box looks identical.
Check the documentation for the actual fields being searched. For a controlled comparison, explicitly specify comparable fields where appropriate rather than relying on defaults whose behavior differs between platforms.
Proximity operators are particularly easy to mistranslate
Databases use different syntax for requiring words to appear near one another. The operators themselves differ, and so can their treatment of word order, maximum distance, punctuation, and field boundaries.
A proximity expression copied literally from another platform may fail, behave differently, or be interpreted as ordinary text.
Translate the intended relationship, such as "these terms should occur within three words of each other in either order," rather than translating only the visible characters of the original query.
Truncation and wildcard behavior can differ too
A truncation symbol may represent any number of characters in one system but behave differently in another. Wildcards can represent optional or single characters, and some platforms impose restrictions on where or how they may be used.
These details matter because a small syntax difference can produce a large retrieval difference. A term intended to capture a controlled family of word variants may become either much narrower or much broader after translation.
Phrase searching may not behave identically
Quotation marks and other phrase-searching mechanisms are not universally interpreted in the same way. Some systems process phrases differently, and some database interfaces apply automatic transformations to unquoted terms.
If a count changes dramatically around a multiword expression, test the expression independently and inspect how the database reports or interprets the query.
Filters and limits may not be equivalent
Applying an apparently similar filter in two databases does not guarantee equivalent selection. Publication types, study designs, age groups, document types, languages, and other metadata may be indexed differently.
Cochrane cautions that search filters should be used with attention to the database and context. It also notes that filters appropriate for one resource may be unnecessary or inappropriate in another. For example, randomized-trial and human filters used for bibliographic databases should not simply be applied to CENTRAL.
When counts diverge sharply, temporarily remove filters and limits. If the discrepancy contracts, investigate how those restrictions operate in each database.
Compare concept blocks rather than only final totals
The final result count tells you that two searches differ, but not why.
A much more useful diagnostic approach is to compare corresponding concept blocks separately:
| Search component |
Database A |
Database B |
Question to investigate |
| Population concept |
42,000 |
47,000 |
Are coverage and terminology reasonably comparable? |
| Intervention concept |
18,000 |
71,000 |
Which subject heading or free-text term causes the divergence? |
| Study-design filter |
3,500,000 |
2,100,000 |
Are equivalent validated filters being used? |
| Final combination |
1,240 |
4,860 |
Does the intervention difference explain most of the final gap? |
These counts are hypothetical, but the diagnostic principle is important. Find the first point at which the strategies diverge substantially, then investigate that component.
More results do not mean better coverage
A database retrieving three times as many records has not necessarily found three times as much relevant evidence. It may contain broader subject coverage, more duplicate-like records, different document types, or a noisier translation of one concept.
Likewise, the smaller result set is not automatically more precise or more accurate.
Judge the searches through relevant retrieval, known-paper testing, inspection of unique records, and the appropriateness of the translated logic rather than by treating result count as a quality score.
Watch Out
Do not "correct" a database search merely because its result count differs from another database. Forcing counts to converge can distort a correctly translated strategy. Explain the difference before deciding that anything needs fixing.
06 · What This Means for You
Troubleshoot the Translation, Not the Number
When database counts differ dramatically, work from the inside out. Compare concept blocks, then individual terms, subject headings, fields, proximity expressions, and filters. The first substantial divergence usually tells you where to investigate.
A simple diagnostic framework
If corresponding concept blocks behave similarly but final totals differ moderately
Database coverage and indexing differences may plausibly explain much of the variation.
If one concept block differs dramatically
Inspect its individual free-text terms, subject headings, fields, and Boolean construction.
If a controlled-vocabulary line causes the discrepancy
Compare scope, hierarchy, explosion, and indexing rather than assuming the headings are direct equivalents.
If a free-text line behaves unexpectedly
Check fields, phrase processing, truncation, wildcards, automatic query processing, and proximity syntax.
If the strategy works in PubMed but the translated version behaves strangely elsewhere
If the differences remain difficult to explain after line-by-line testing, external review may be worthwhile. Search translation is one of the points at which a research librarian or information specialist can identify database-specific problems that are easy to overlook.
07 · A Quick Checklist
What to Compare When Database Result Counts Differ
Before deciding that one search is wrong, check:
Confirm that each database actually covers the relevant disciplines, journals, publication periods, and document types.
Compare corresponding concept blocks separately rather than comparing only the final result totals.
Map controlled-vocabulary terms independently in each database and inspect their scope and hierarchy.
Verify whether subject headings are exploded and what narrower concepts the explosion includes.
Check which fields each free-text expression actually searches.
Translate proximity, truncation, wildcard, and phrase syntax according to the destination platform.
Check whether filters and limits are genuinely comparable and appropriate in both resources.
Test known relevant papers and inspect unique relevant retrieval rather than judging quality from counts alone.