03 · What You Need to Know
Identifiability can emerge from combinations
Indirect identifiers do not have to identify someone by themselves
A direct identifier points relatively straightforwardly to a person. A name, for example, may directly identify someone in the relevant context.
Indirect identifiers work differently. Information Commissioner's Office guidance explains that information can identify someone indirectly when it is combined with other information. It gives a combination such as age, occupation, and place of residence as an example of criteria that may allow a person to be singled out.
UK Data Service materials similarly describe indirect identifiers as information that may uniquely identify people in combination and give examples including gender, age, region, occupation, and income.
The practical implication is important: a variable does not have to be identifying alone to contribute to identification.
Think of identification as narrowing a population
One way to understand combination risk is to imagine each detail narrowing the set of possible people.
Location
Begin with everyone living in a municipality.
Occupation
Narrow the group to university professors in that municipality.
Specialty
Narrow it again to professors in one uncommon academic discipline.
Age
Add a narrow age range.
Distinctive event
Add a recently publicized award or appointment, and perhaps only one plausible person remains.
No individual detail had to contain the person's name. Identification emerged through progressive narrowing.
This is often called jigsaw identification
The metaphor is useful because each piece reveals only part of the picture. Once enough pieces are assembled, the identity becomes apparent.
UK Data Service guidance on qualitative text explicitly notes that disclosure risk frequently arises from combinations of contextual detail and refers to this as jigsaw identification. Examples include rare occupations in small communities, distinctive career trajectories, detailed timelines, highly specific locations, unique personal experiences, and information about identifiable third parties.
The concept applies beyond interview transcripts. Demographic tables, survey microdata, case descriptions, fieldnotes, photographs, geographic information, administrative datasets, and mixed datasets can all contain combinations that distinguish participants.
Singling someone out can matter even before you know their name
Identification is sometimes misunderstood as requiring a person's name. Current ICO anonymisation guidance uses singling out and linkability as key indicators of identifiability.
Singling out means being able to isolate information relating to one person from information relating to others. The ICO notes that even when someone does not intend to act on that information, the ability to single the person out can mean they remain identifiable.
This distinction matters for research datasets. A record might uniquely describe "the 47-year-old neurosurgeon in Municipality X" even if the dataset never states the person's name. Other information may then make linking that record to a named individual straightforward.
Linkability allows one dataset to supply the missing pieces
The combination does not have to exist entirely inside your research dataset.
A third party may combine research information with staff directories, professional profiles, institutional websites, news articles, public records, social media, published biographies, or another dataset. ICO guidance specifically warns that information which does not directly identify someone may become identifying when combined with information held elsewhere.
Researchers should therefore ask not only what their dataset contains but what information is reasonably available to the people likely to receive it.
Population size changes how identifying a combination is
The same characteristics can present very different risks in different populations.
"Female, 45, teacher" may describe many thousands of people nationally. In a study involving six employees from one small school, the same characteristics may point to a single participant.
ICO guidance gives a similar contextual example: a person's year of birth may distinguish them within one small group but not within a larger population. The risk of singling out therefore depends on the context in which the information appears.
This is why fixed lists of "safe" demographic variables are unreliable. Risk depends partly on how common or rare the combination is in the relevant population.
Precision increases the narrowing power of a detail
Specific information generally narrows a population more than broad information.
More specific
Age 43, exact municipality, exact job title, exact date of an event.
Less specific
Age 40–49, broader region, occupational category, approximate period.
Generalization is therefore one common disclosure-control technique. Current ICO materials describe generalization as aggregating information to a higher level of abstraction, such as age groups or geographic regions.
Reducing precision can make more people fit the same description. The trade-off is that it may also reduce analytical usefulness.
Rare characteristics deserve particular attention
A characteristic shared by almost everyone in a sample may contribute little to identification. A characteristic possessed by only one participant can be far more revealing.
Rare occupations, uncommon diagnoses, unusual family structures, unique professional histories, distinctive awards, rare combinations of qualifications, or extraordinary events can act as powerful narrowing clues.
Qualitative material is especially challenging because participants often explain their experiences through precisely these distinctive details. UK Data Service guidance highlights rare occupations, distinctive career trajectories, and unique personal experiences as examples of contextual information that can contribute to identification.
The combination can span an entire publication
Researchers should not assess combination risk only within a single table or quotation.
A methods section might reveal the institution. A participant table provides age ranges and occupations. One quotation identifies a professional specialty. Another quotation under the same participant code describes a recent event. Together, these pieces may create a much more identifiable profile than any section does independently.
This is one reason an apparently anonymous quotation can still identify a participant. The quotation may supply only the final piece of a puzzle built elsewhere in the publication.
Participants who know one another create an especially difficult environment
Insiders already possess pieces of the puzzle.
In a workplace study, colleagues may know one another's ages, roles, family situations, recent promotions, conflicts, or notable experiences. In a small community, participants may recognize events or relationships immediately.
This means a combination that looks obscure to the researcher may be transparent to another participant. Studies involving interconnected participants therefore require particular attention to confidentiality when participants already know one another.
Not every imaginable combination makes data identifiable
Combination risk should not be interpreted as meaning that any theoretical possibility of re-identification makes anonymisation impossible.
Under current UK guidance, the assessment asks whether identification means are reasonably likely to be used and considers objective factors such as available information, technology, cost, and time. A very remote hypothetical possibility is not treated the same way as a practical route to identification.
Other jurisdictions use their own standards, so researchers should apply the rules relevant to their study. The general methodological lesson remains useful: consider realistic combinations and realistic external information rather than either ignoring linkage risk or imagining every conceivable adversary.
Watch Out
Do not approve age, occupation, location, institution, and event history separately and then assume the dataset is safe. The confidentiality question concerns the profile those details create together.
07 · A Quick Checklist
Before deciding that combined participant details are safe
Check combinations involving:
Age or narrow age ranges combined with occupation, role, or seniority.
Specific locations combined with demographic or professional characteristics.
Rare occupations, specialties, diagnoses, qualifications, or institutional positions.
Detailed timelines, appointments, awards, incidents, or distinctive life events.
Information spread across methods sections, demographic tables, quotations, case descriptions, and appendices.
Repeated participant codes that allow readers to combine details across several excerpts.
Information available from realistic external sources that could be linked with the research data.
Knowledge likely to be held by colleagues, relatives, other participants, or community members familiar with the setting.
Whether reducing precision or removing one unnecessary clue would meaningfully reduce singling-out or linkage risk.