Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Several Harmless Details Become Identifying When They Appear Together?

A participant may be identifiable even when no single detail reveals who they are. Several ordinary characteristics can combine into a distinctive profile that points to one person.

301
When Combined Details Identify Participants Guide 301 of 398
01 · The Question

Can information become identifying only after you put the pieces together?

A participant is 52 years old. That alone may reveal little. They are a school principal. Still not necessarily identifying. They live in a particular municipality. Again, perhaps many people fit.

Now combine the details: a 52-year-old school principal in that municipality who recently completed a doctorate overseas and received a particular national award.

Suddenly, a collection of ordinary characteristics may describe one person remarkably well. Research confidentiality can fail not because one obvious identifier was left behind, but because several weaker clues work together.

02 · The Short Answer

Several weak identifiers can form one strong identifying profile

In Brief

Yes. Details such as age, occupation, location, gender, institution, education, family structure, chronology, or distinctive experiences may not identify a participant individually but can become identifying when combined with one another or linked with information available elsewhere.

This is commonly described as indirect identification, jigsaw identification, or, in some technical contexts, identification through combinations of quasi-identifiers. Researchers should therefore assess combinations of information rather than deciding that each variable is safe in isolation.

03 · What You Need to Know

Identifiability can emerge from combinations

Indirect identifiers do not have to identify someone by themselves

A direct identifier points relatively straightforwardly to a person. A name, for example, may directly identify someone in the relevant context.

Indirect identifiers work differently. Information Commissioner's Office guidance explains that information can identify someone indirectly when it is combined with other information. It gives a combination such as age, occupation, and place of residence as an example of criteria that may allow a person to be singled out.

UK Data Service materials similarly describe indirect identifiers as information that may uniquely identify people in combination and give examples including gender, age, region, occupation, and income.

The practical implication is important: a variable does not have to be identifying alone to contribute to identification.

Think of identification as narrowing a population

One way to understand combination risk is to imagine each detail narrowing the set of possible people.

Location Begin with everyone living in a municipality.
Occupation Narrow the group to university professors in that municipality.
Specialty Narrow it again to professors in one uncommon academic discipline.
Age Add a narrow age range.
Distinctive event Add a recently publicized award or appointment, and perhaps only one plausible person remains.

No individual detail had to contain the person's name. Identification emerged through progressive narrowing.

This is often called jigsaw identification

The metaphor is useful because each piece reveals only part of the picture. Once enough pieces are assembled, the identity becomes apparent.

UK Data Service guidance on qualitative text explicitly notes that disclosure risk frequently arises from combinations of contextual detail and refers to this as jigsaw identification. Examples include rare occupations in small communities, distinctive career trajectories, detailed timelines, highly specific locations, unique personal experiences, and information about identifiable third parties.

The concept applies beyond interview transcripts. Demographic tables, survey microdata, case descriptions, fieldnotes, photographs, geographic information, administrative datasets, and mixed datasets can all contain combinations that distinguish participants.

Singling someone out can matter even before you know their name

Identification is sometimes misunderstood as requiring a person's name. Current ICO anonymisation guidance uses singling out and linkability as key indicators of identifiability.

Singling out means being able to isolate information relating to one person from information relating to others. The ICO notes that even when someone does not intend to act on that information, the ability to single the person out can mean they remain identifiable.

This distinction matters for research datasets. A record might uniquely describe "the 47-year-old neurosurgeon in Municipality X" even if the dataset never states the person's name. Other information may then make linking that record to a named individual straightforward.

Linkability allows one dataset to supply the missing pieces

The combination does not have to exist entirely inside your research dataset.

A third party may combine research information with staff directories, professional profiles, institutional websites, news articles, public records, social media, published biographies, or another dataset. ICO guidance specifically warns that information which does not directly identify someone may become identifying when combined with information held elsewhere.

Researchers should therefore ask not only what their dataset contains but what information is reasonably available to the people likely to receive it.

Population size changes how identifying a combination is

The same characteristics can present very different risks in different populations.

"Female, 45, teacher" may describe many thousands of people nationally. In a study involving six employees from one small school, the same characteristics may point to a single participant.

ICO guidance gives a similar contextual example: a person's year of birth may distinguish them within one small group but not within a larger population. The risk of singling out therefore depends on the context in which the information appears.

This is why fixed lists of "safe" demographic variables are unreliable. Risk depends partly on how common or rare the combination is in the relevant population.

Precision increases the narrowing power of a detail

Specific information generally narrows a population more than broad information.

More specific Age 43, exact municipality, exact job title, exact date of an event.
Less specific Age 40–49, broader region, occupational category, approximate period.

Generalization is therefore one common disclosure-control technique. Current ICO materials describe generalization as aggregating information to a higher level of abstraction, such as age groups or geographic regions.

Reducing precision can make more people fit the same description. The trade-off is that it may also reduce analytical usefulness.

Rare characteristics deserve particular attention

A characteristic shared by almost everyone in a sample may contribute little to identification. A characteristic possessed by only one participant can be far more revealing.

Rare occupations, uncommon diagnoses, unusual family structures, unique professional histories, distinctive awards, rare combinations of qualifications, or extraordinary events can act as powerful narrowing clues.

Qualitative material is especially challenging because participants often explain their experiences through precisely these distinctive details. UK Data Service guidance highlights rare occupations, distinctive career trajectories, and unique personal experiences as examples of contextual information that can contribute to identification.

The combination can span an entire publication

Researchers should not assess combination risk only within a single table or quotation.

A methods section might reveal the institution. A participant table provides age ranges and occupations. One quotation identifies a professional specialty. Another quotation under the same participant code describes a recent event. Together, these pieces may create a much more identifiable profile than any section does independently.

This is one reason an apparently anonymous quotation can still identify a participant. The quotation may supply only the final piece of a puzzle built elsewhere in the publication.

Participants who know one another create an especially difficult environment

Insiders already possess pieces of the puzzle.

In a workplace study, colleagues may know one another's ages, roles, family situations, recent promotions, conflicts, or notable experiences. In a small community, participants may recognize events or relationships immediately.

This means a combination that looks obscure to the researcher may be transparent to another participant. Studies involving interconnected participants therefore require particular attention to confidentiality when participants already know one another.

Not every imaginable combination makes data identifiable

Combination risk should not be interpreted as meaning that any theoretical possibility of re-identification makes anonymisation impossible.

Under current UK guidance, the assessment asks whether identification means are reasonably likely to be used and considers objective factors such as available information, technology, cost, and time. A very remote hypothetical possibility is not treated the same way as a practical route to identification.

Other jurisdictions use their own standards, so researchers should apply the rules relevant to their study. The general methodological lesson remains useful: consider realistic combinations and realistic external information rather than either ignoring linkage risk or imagining every conceivable adversary.

Watch Out

Do not approve age, occupation, location, institution, and event history separately and then assume the dataset is safe. The confidentiality question concerns the profile those details create together.

04 · A Practical Example

Five ordinary details narrow a sample to one person

Hypothetical Example

A study reports a participant profile without giving a name

A study of university faculty describes one participant as a woman aged 40–49, a professor at a university in a particular small city, working in a rare academic specialty, who returned from an overseas fellowship the previous year.

Gender and age Many people fit.
City and institution type The possible population becomes smaller.
Rare academic specialty Only a few plausible individuals may remain.
Recent overseas fellowship An institutional news page or professional profile may point to one person.
Result No individual field contains a name, but the combination may allow the participant to be singled out and linked to publicly available information.

The appropriate response might be to broaden the geographic description, remove the fellowship detail, generalize the specialty, reduce demographic precision, or use some combination of these approaches. Which detail should change depends on what the analysis actually needs.

05 · What Researchers Often Get Wrong

Common mistakes when assessing combinations of identifiers

Misconception

None of these variables identifies anyone by itself, so the dataset is anonymous

Indirect identification often arises from combinations. Several non-unique characteristics can collectively distinguish one participant from everyone else.

Misconception

Only information inside my dataset matters

External information can supply the missing link. Institutional websites, professional profiles, public records, news reports, social media, and other datasets may make an otherwise obscure profile identifiable.

Misconception

If I remove names, combination risk disappears

Names are direct identifiers. Removing them does not prevent indirect characteristics such as age, occupation, location, or distinctive experiences from working together.

Misconception

The same demographic combination has the same risk everywhere

A profile that describes thousands of people nationally may describe one person within a small organization, village, clinic, or specialist professional community. Population and audience matter.

Misconception

I only need to check each table and quotation separately

Identification can emerge across an entire article, report, thesis, or dataset. Methods, demographic tables, participant codes, quotations, and contextual descriptions can supply different pieces of the same profile.

Misconception

Any theoretical combination means the data can never be anonymous

Identifiability standards are contextual. For example, current UK guidance considers means reasonably likely to be used rather than every remote hypothetical possibility. Apply the appropriate standard for your jurisdiction and disclosure context.

06 · What This Means for You

Review profiles, not just variables

When checking a dataset or publication for confidentiality, stop asking only whether each field is identifying. Look at the participant profile created when fields are combined.

A simple combination-risk framework

If several details describe the same participant
Evaluate them together rather than approving each one independently.
If a combination is rare in the relevant population
Consider reducing precision, suppressing an unnecessary detail, or otherwise reducing the ability to single the participant out.
If external information could plausibly complete the identification
Include that realistic linkage in the disclosure-risk assessment.
If participants or readers know the setting well
Assess what those insiders already know rather than evaluating identifiability only from a stranger's perspective.
If reducing combination risk would remove analytically important information
Consider which details can be generalized, whether another form of reporting is possible, or whether stronger access controls are more appropriate.

A useful final test is to read the participant profile as though you were trying to identify the person. You do not need to become a privacy attacker for the afternoon, but the change in perspective often reveals combinations that are easy to miss when reviewing variables one column at a time.

07 · A Quick Checklist

Before deciding that combined participant details are safe

Check combinations involving:
Age or narrow age ranges combined with occupation, role, or seniority.
Specific locations combined with demographic or professional characteristics.
Rare occupations, specialties, diagnoses, qualifications, or institutional positions.
Detailed timelines, appointments, awards, incidents, or distinctive life events.
Information spread across methods sections, demographic tables, quotations, case descriptions, and appendices.
Repeated participant codes that allow readers to combine details across several excerpts.
Information available from realistic external sources that could be linked with the research data.
Knowledge likely to be held by colleagues, relatives, other participants, or community members familiar with the setting.
Whether reducing precision or removing one unnecessary clue would meaningfully reduce singling-out or linkage risk.
08 · Frequently Asked Questions

Questions about jigsaw identification and combined details

What is jigsaw identification?

Jigsaw identification occurs when separate pieces of information can be combined to reveal or infer a person's identity. In qualitative research, these pieces may include occupation, location, chronology, relationships, and distinctive experiences.

What is an indirect identifier?

An indirect identifier is information that may contribute to identifying someone when combined with other information. Examples can include age, occupation, location, education, income, or other characteristics depending on context.

Are indirect identifiers the same as quasi-identifiers?

The terms are often used in closely related ways, particularly in disclosure-control and privacy literature. ICO guidance describes indirect identifiers as pieces or combinations of information used to identify a person and notes that they are also sometimes called quasi-identifiers.

How many details does it take to identify someone?

There is no universal number. Risk depends on how specific and rare the characteristics are, the size of the relevant population, what external information exists, and who receives the data. One highly distinctive characteristic may be more revealing than several common ones.

Can public information be combined with research data to identify participants?

Yes. Professional profiles, institutional websites, news reports, public records, social media, and other sources may sometimes provide information that can be linked with research data. The relevant assessment should consider realistic sources available to likely recipients.

Does grouping ages into ranges solve the problem?

It can reduce precision and therefore identification risk, but it does not guarantee anonymity. Age range may still combine with location, occupation, institutional role, or other characteristics to single someone out.

Is jigsaw identification mainly a qualitative research problem?

No. Rich qualitative narratives make the problem especially visible, but combinations of indirect identifiers can also affect survey microdata, administrative records, health datasets, geographic data, photographs, and other research materials.

09 · The Bottom Line

Harmless pieces can form an identifying picture

The Bottom Line

Several details that do not identify a participant individually can become identifying when combined, particularly when the resulting profile is rare or can be linked with information available elsewhere.

Assess participant profiles rather than isolated variables. Consider population size, precision, rarity, external information, and what likely readers already know. In confidentiality work, the troublesome identifier is sometimes not one piece of data at all. It is the picture the pieces make together.

10 · Sources and Further Reading

Authoritative guidance on indirect and combined identification

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes