03 · What You Need to Know
Small Samples Increase Risk When They Make People Easier to Single Out
The Problem Is Uniqueness, Not the Number Five, Ten, or Twenty
Researchers sometimes look for a numerical rule: below what sample size do data become identifiable?
There is no universal threshold.
A sample of five participants drawn anonymously from millions of people may reveal very little about who they are. A sample of 100 people consisting of almost every member of one highly visible professional group may still contain distinctive individuals.
The relevant question is whether information allows a person to be singled out or linked to other information. The ICO identifies singling out and linkability as key indicators of identifiability and emphasizes that the risk depends on context and the richness of the data.
Small Samples Shrink the Set of Possible People
Suppose a national survey reports that one participant is a 52-year-old male professor. Thousands of people may match.
Now suppose the study involves six professors from one small department. If only one is male and approximately 52, the same information becomes far more revealing.
Nothing about "52" or "male" changed intrinsically. The comparison population changed.
This illustrates why identifiability is contextual. A year of birth may distinguish someone within a family while doing little to distinguish them within a large school cohort, an example used in current ICO guidance.
A Small Sample Can Turn Indirect Identifiers Into Near-Direct Identifiers
Age, occupation, gender, academic rank, location, years of service, educational background, and similar variables are often indirect identifiers. In large populations, many people may share each value.
Within a small sample, combinations become sparse.
For example:
- female;
- full professor;
- engineering;
- over 60;
- more than 30 years at the institution.
If only one person fits that description, the combination effectively singles her out.
This is why researchers should examine how demographic details interact rather than evaluating each characteristic independently.
The Underlying Population Can Matter More Than the Sample
A study may recruit only ten participants but draw them from a population of thousands with no public sampling frame. Another may recruit ten of the twelve people holding a particular position in one organisation.
The second situation can create substantially more recognition risk because readers may already know the small universe from which participants were drawn.
Researchers should therefore distinguish the sample size from the size and visibility of the source population.
Small sample
The research includes relatively few participants.
Small identifiable population
The possible people from whom participants could have been drawn are few, distinctive, or already known to likely readers.
The two often coincide, but they are not the same problem.
Sampling Information Can Reveal More Than Researchers Expect
A paper may carefully remove participant names while describing recruitment in enough detail to reconstruct the participant pool.
Statements such as "all four female deans at University X participated" effectively identify membership in the sample even if subsequent quotations use participant codes.
Likewise, saying that "three of the institution's four vice presidents participated" sharply narrows the possible speakers whenever a quotation is attributed to a vice president.
Confidentiality review should therefore include the Methods section, sample table, appendices, supplementary files, and other contextual information, not only quotations and raw data.
Quotations Can Be More Identifying Than Demographic Tables
Qualitative quotations contain vocabulary, events, relationships, opinions, and biographical information that may be recognisable to people familiar with participants.
A quotation such as "when I became dean immediately after the 2024 merger" may identify someone more readily than an age or gender variable.
HHS advisory materials have recognised that qualitative recordings may be identifiable depending on circumstances such as a limited sample and unique voice characteristics.
OHRP's exploratory work on qualitative research also notes risks from sensitive or identifiable participant and third-party information and discusses measures such as redaction, site pseudonyms, and attention to dissemination.
Participants May Recognise Themselves and One Another
A small-sample report may permit participants to identify other participants even when outside readers cannot.
This can be particularly important in research involving illegal or socially sensitive behaviour. SACHRP has noted that in small studies, returning general results can sometimes create a possibility that participants identify themselves or other participants.
The likely audience therefore matters. Researchers should consider recognition by participants, colleagues, employers, community members, and other insiders, not merely an anonymous public reader.
Reporting Percentages Can Accidentally Reveal Individuals
Small samples create a less obvious disclosure problem: percentages can look aggregated while still representing very few people.
If a study contains eight participants, 12.5% represents one participant. Reporting that "12.5% of participants disclosed condition X" may therefore communicate that exactly one person did so.
If the paper also provides demographic breakdowns or quotations, readers may be able to determine which participant that was.
Aggregation is useful only when the resulting groups are sufficiently non-distinctive for the intended release context.
Cross-Tabulation Can Create Tiny Cells
A table may look safe because no names appear. But intersecting categories can create cells containing one or two people.
For example:
| Rank |
Male |
Female |
Total |
| Professor |
4 |
1 |
5 |
| Associate Professor |
2 |
3 |
5 |
| Assistant Professor |
0 |
4 |
4 |
If readers know the institution has only one female professor in the relevant population, a result reported specifically for that cell may effectively become participant-level information.
There is no universal minimum cell size that guarantees anonymity across all research contexts. Statistical agencies, repositories, institutions, and regulatory frameworks may impose specific suppression rules, but researchers should use the rule governing their dataset rather than inventing one.
Removing Demographics Can Reduce Risk but Also Damage the Research
The solution is not necessarily to delete every participant characteristic.
Demographics may be necessary to interpret transferability, describe inequities, examine subgroup differences, or understand the phenomenon being studied. Qualitative research may depend on role and context for interpretation.
Researchers need to balance disclosure risk with scientific meaning. Options may include broader categories, suppression of particular combinations, less precise site descriptions, controlled access, modified attribution of quotations, or withholding details only where they create disproportionate identification risk.
The objective is not to make every participant indistinguishable from every human being on Earth. It is to manage realistic identification risk without quietly deleting the variables that made the study worth conducting.
Combining Sections of a Paper Can Reconstruct Identity
Disclosure risk often emerges across the article rather than within one table.
The Methods section identifies the institutions. Table 1 provides age, gender, and rank. The Results section attributes quotations by participant code. An appendix lists years of experience. Each element may appear acceptable alone.
Together, they may provide enough information to reconstruct identities.
This is essentially the same linkage problem that occurs when multiple datasets are combined. In publication, the "datasets" may simply be different tables, quotations, and contextual descriptions within the same paper.
Watch Out
Do not assess disclosure risk one table or quotation at a time. A reader sees the entire paper and may also know the study setting. Details that appear harmless separately can become identifying when combined.
Small Samples Do Not Automatically Make Confidentiality Impossible
Small-sample research is essential in case studies, rare populations, specialist professions, qualitative inquiry, pilot studies, organisational research, and many other designs.
The appropriate response is not to declare such research inherently non-confidential. It is to recognise that confidentiality may require more contextual judgment than simply removing names.
| Factor |
Lower Identification Concern |
Higher Identification Concern |
| Source population |
Large and not readily known |
Small, bounded, or publicly known |
| Participant characteristics |
Common, broad categories |
Rare or unique combinations |
| Geography or institution |
Broadly described |
Precisely identified small setting |
| Quotations |
Non-distinctive content |
Unique events, language, or biographical clues |
| Reporting |
Groups remain sufficiently non-distinctive |
Single-person or tiny cells |
| Audience knowledge |
Little contextual knowledge |
Colleagues, community members, or other insiders |