01 · The Question
What can identify a research participant besides their name?
Researchers often begin de-identification by removing obvious information such as names, email addresses, telephone numbers, or identification numbers. That is sensible, but it can create a false sense of completion.
A participant may still be identifiable through their face, voice, location, occupation, online identifier, distinctive experience, social relationships, or other characteristics. Sometimes no single detail identifies them. The problem emerges only when several details are combined.
The practical question is therefore broader than "Did I remove the direct identifiers?" You need to consider what information remains and whether it can distinguish or point back to a particular person.
03 · What You Need to Know
Identifiability depends on context, not just obvious identifiers
Direct identifiers are only the easiest cases
Some identifiers are obvious. A participant's full name, personal email address, telephone number, or another unique identification number may point directly to one person.
Research data can also identify people indirectly. The UK GDPR definition of personal data, for example, expressly includes direct or indirect identification and refers to identifiers such as names, identification numbers, location data, and online identifiers, as well as factors relating to a person's physical, physiological, genetic, mental, economic, cultural, or social identity.
The U.S. Common Rule uses different terminology. It defines identifiable private information as private information for which a participant's identity is or may readily be ascertained by the investigator or associated with the information. OHRP training materials emphasize that this determination is contextual rather than based on a fixed list of identifiers.
These frameworks are not interchangeable legal tests, but they illustrate the same practical problem for researchers: deleting a name does not settle whether the remaining data identify someone.
A face can identify someone without accompanying text
A clear photograph or video may permit people who know the participant to recognize them immediately. Other visible features can also contribute to identification, including distinctive tattoos, scars, clothing, uniforms, assistive devices, name badges, or surroundings.
Researchers should be precise with the term biometric data. A photograph of a face is not automatically biometric data under every legal definition simply because the person can be recognized. Under UK GDPR terminology, biometric data require specific technical processing of physical, physiological, or behavioural characteristics that allows or confirms unique identification. Facial-recognition processing is a clear example.
The narrower biometric classification does not change the basic confidentiality issue. A normal photograph can still be identifiable personal information because a person is visible and recognizable.
A voice can identify someone even without a spoken name
Voice is another obvious example of information that can survive conventional de-identification. A person familiar with a participant may recognize them from the sound of their voice, while accent, dialect, speech patterns, or what they say can provide additional clues.
For this reason, an audio recording can remain identifiable without a participant ever saying their name. Replacing the filename with a participant code does not remove those characteristics from the recording.
Location can become identifying when it narrows the possibilities
Location information varies enormously in specificity. "Southeast Asia" generally reveals much less than a street address. Between those extremes are countries, provinces, cities, villages, workplaces, schools, clinics, GPS coordinates, travel routes, and repeated location patterns.
The importance of a location also depends on the population. A job title combined with "Manila" might describe many people, while the same occupation combined with a very small municipality could narrow the possibilities dramatically. UK data protection guidance expressly recognizes location data as a potential identifier, while qualitative-data guidance recommends considering highly specific geographic references during anonymisation.
Researchers should therefore ask what geographic precision is analytically necessary. Generalizing a location from a specific facility to a broader region may sometimes preserve the relevant analytical meaning while reducing disclosure risk, although this depends on the study.
Occupation and institutional affiliation can identify people in small populations
Job information is easy to underestimate because occupations are shared by many people. Context changes the calculation.
"University professor" is broad. "The only professor of a particular specialty at a named institution" may point toward one person. The same can occur with senior positions, rare occupations, unusual qualifications, specific departments, or combinations of employer and role.
UK Data Service guidance on qualitative text specifically warns that a rare occupation in a small community, a distinctive career trajectory, detailed timelines, specific locations, and unique personal experiences can permit identification through narrative context.
Distinctive events and life histories can function like identifiers
Qualitative researchers often collect detailed narratives precisely because specificity matters. Unfortunately, specificity can also disclose identity.
A participant might describe being the first person from their organization to receive a particular award, surviving a widely reported accident, leading a public controversy, occupying an unusual professional role, or experiencing a rare sequence of events. Removing the participant's name does little if those events can be searched or are already known within the relevant community.
This is why anonymisation of qualitative material requires judgment rather than simple find-and-replace. Altering too much can damage analytical meaning, while retaining too much can expose the participant. When those goals conflict, researchers need to consider what to do when removing identifying details changes the meaning of qualitative data.
Online identifiers and digital traces can also distinguish people
Digital research introduces additional routes to identification. Current Information Commissioner's Office guidance identifies IP addresses, cookie identifiers, MAC addresses, advertising IDs, account handles, and device fingerprints as examples of online identifiers that may constitute personal data depending on context.
A username does not have to reveal a legal name to distinguish one person from another. The same applies to persistent digital identifiers that permit records or activities to be linked over time.
The combination can matter more than any individual detail
Consider a dataset containing age, occupation, municipality, educational history, and a distinctive life event. None of those variables necessarily identifies the participant alone. Together, they may describe only one plausible person.
The Information Commissioner's Office describes this as linkability and notes that the combination of separate information can create a mosaic or jigsaw effect. HHS advisory material has likewise recognized that multiple data points that do not individually identify someone may, in aggregate, render information readily identifiable.
Identifiability depends partly on who receives the data
The same information can present different identification risks in different contexts. A quotation may mean nothing to the general public but be recognizable to the participant's coworkers. A photograph of an office may reveal an employer to local staff. A rare diagnosis combined with age and location may narrow a patient population substantially.
This does not mean researchers must protect against every theoretically imaginable identification scenario. Under current UK anonymisation guidance, for example, the assessment considers means reasonably likely to be used, taking into account factors such as available information, technology, time, and cost. Other jurisdictions use their own legal standards.
The broader research lesson is that identifiability should be evaluated in the actual disclosure environment rather than imagined in a vacuum.
Pseudonymisation reduces risk but does not necessarily remove identifiability
Replacing names with participant codes can be an important safeguard. It separates obvious identifying information from the working dataset and can reduce unnecessary exposure.
But if additional information exists that reconnects the code to the person, the data have not necessarily become anonymous. Likewise, even without the code key, indirect information in the dataset may still permit identification.
Researchers should therefore distinguish between reducing identification risk and eliminating it. That distinction affects how identifiable research materials should be collected, accessed, shared, retained, and disposed of.
07 · A Quick Checklist
Before deciding that research data are non-identifiable
Check whether participants could be distinguished through:
Names, contact details, identification numbers, account names, or other direct identifiers.
Recognizable faces, voices, distinctive physical features, or other observable characteristics.
Specific locations, workplaces, schools, institutions, or geographic patterns.
Rare occupations, senior roles, unusual qualifications, or distinctive career histories.
Unique events, relationships, timelines, experiences, or narrative details.
Online identifiers, usernames, device information, or other persistent digital traces where relevant.
Combinations of details that are substantially more identifying together than separately.
External information reasonably available to the people who will access or receive the data.
A code key or other linkage mechanism that can reconnect pseudonymised records to participants.