Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Voice, Face, Location, or Other Characteristics Make Research Data Identifiable?

Names are only one route to identification. Voice, face, location, occupation, distinctive experiences, and combinations of contextual details can also reveal who a research participant is.

293
What Can Make Research Data Identifiable? Guide 293 of 398
01 · The Question

What can identify a research participant besides their name?

Researchers often begin de-identification by removing obvious information such as names, email addresses, telephone numbers, or identification numbers. That is sensible, but it can create a false sense of completion.

A participant may still be identifiable through their face, voice, location, occupation, online identifier, distinctive experience, social relationships, or other characteristics. Sometimes no single detail identifies them. The problem emerges only when several details are combined.

The practical question is therefore broader than "Did I remove the direct identifiers?" You need to consider what information remains and whether it can distinguish or point back to a particular person.

02 · The Short Answer

Almost any sufficiently distinctive information can contribute to identification

In Brief

Yes. Voice, face, location, physical characteristics, occupation, institutional affiliation, online identifiers, distinctive events, relationships, and other contextual information can make research data identifiable, either by themselves or when combined with other information.

There is no universal checklist that determines identifiability in every study. The relevant question is whether a person can reasonably be distinguished, recognized, or linked to the data in the circumstances in which those data are held or disclosed.

03 · What You Need to Know

Identifiability depends on context, not just obvious identifiers

Direct identifiers are only the easiest cases

Some identifiers are obvious. A participant's full name, personal email address, telephone number, or another unique identification number may point directly to one person.

Research data can also identify people indirectly. The UK GDPR definition of personal data, for example, expressly includes direct or indirect identification and refers to identifiers such as names, identification numbers, location data, and online identifiers, as well as factors relating to a person's physical, physiological, genetic, mental, economic, cultural, or social identity.

The U.S. Common Rule uses different terminology. It defines identifiable private information as private information for which a participant's identity is or may readily be ascertained by the investigator or associated with the information. OHRP training materials emphasize that this determination is contextual rather than based on a fixed list of identifiers.

These frameworks are not interchangeable legal tests, but they illustrate the same practical problem for researchers: deleting a name does not settle whether the remaining data identify someone.

A face can identify someone without accompanying text

A clear photograph or video may permit people who know the participant to recognize them immediately. Other visible features can also contribute to identification, including distinctive tattoos, scars, clothing, uniforms, assistive devices, name badges, or surroundings.

Researchers should be precise with the term biometric data. A photograph of a face is not automatically biometric data under every legal definition simply because the person can be recognized. Under UK GDPR terminology, biometric data require specific technical processing of physical, physiological, or behavioural characteristics that allows or confirms unique identification. Facial-recognition processing is a clear example.

The narrower biometric classification does not change the basic confidentiality issue. A normal photograph can still be identifiable personal information because a person is visible and recognizable.

A voice can identify someone even without a spoken name

Voice is another obvious example of information that can survive conventional de-identification. A person familiar with a participant may recognize them from the sound of their voice, while accent, dialect, speech patterns, or what they say can provide additional clues.

For this reason, an audio recording can remain identifiable without a participant ever saying their name. Replacing the filename with a participant code does not remove those characteristics from the recording.

Location can become identifying when it narrows the possibilities

Location information varies enormously in specificity. "Southeast Asia" generally reveals much less than a street address. Between those extremes are countries, provinces, cities, villages, workplaces, schools, clinics, GPS coordinates, travel routes, and repeated location patterns.

The importance of a location also depends on the population. A job title combined with "Manila" might describe many people, while the same occupation combined with a very small municipality could narrow the possibilities dramatically. UK data protection guidance expressly recognizes location data as a potential identifier, while qualitative-data guidance recommends considering highly specific geographic references during anonymisation.

Researchers should therefore ask what geographic precision is analytically necessary. Generalizing a location from a specific facility to a broader region may sometimes preserve the relevant analytical meaning while reducing disclosure risk, although this depends on the study.

Occupation and institutional affiliation can identify people in small populations

Job information is easy to underestimate because occupations are shared by many people. Context changes the calculation.

"University professor" is broad. "The only professor of a particular specialty at a named institution" may point toward one person. The same can occur with senior positions, rare occupations, unusual qualifications, specific departments, or combinations of employer and role.

UK Data Service guidance on qualitative text specifically warns that a rare occupation in a small community, a distinctive career trajectory, detailed timelines, specific locations, and unique personal experiences can permit identification through narrative context.

Distinctive events and life histories can function like identifiers

Qualitative researchers often collect detailed narratives precisely because specificity matters. Unfortunately, specificity can also disclose identity.

A participant might describe being the first person from their organization to receive a particular award, surviving a widely reported accident, leading a public controversy, occupying an unusual professional role, or experiencing a rare sequence of events. Removing the participant's name does little if those events can be searched or are already known within the relevant community.

This is why anonymisation of qualitative material requires judgment rather than simple find-and-replace. Altering too much can damage analytical meaning, while retaining too much can expose the participant. When those goals conflict, researchers need to consider what to do when removing identifying details changes the meaning of qualitative data.

Online identifiers and digital traces can also distinguish people

Digital research introduces additional routes to identification. Current Information Commissioner's Office guidance identifies IP addresses, cookie identifiers, MAC addresses, advertising IDs, account handles, and device fingerprints as examples of online identifiers that may constitute personal data depending on context.

A username does not have to reveal a legal name to distinguish one person from another. The same applies to persistent digital identifiers that permit records or activities to be linked over time.

The combination can matter more than any individual detail

Consider a dataset containing age, occupation, municipality, educational history, and a distinctive life event. None of those variables necessarily identifies the participant alone. Together, they may describe only one plausible person.

The Information Commissioner's Office describes this as linkability and notes that the combination of separate information can create a mosaic or jigsaw effect. HHS advisory material has likewise recognized that multiple data points that do not individually identify someone may, in aggregate, render information readily identifiable.

Watch Out

Do not evaluate indirect identifiers one at a time and conclude that the dataset is safe because each individual detail seems harmless. Several apparently harmless details can become identifying when they appear together.

Identifiability depends partly on who receives the data

The same information can present different identification risks in different contexts. A quotation may mean nothing to the general public but be recognizable to the participant's coworkers. A photograph of an office may reveal an employer to local staff. A rare diagnosis combined with age and location may narrow a patient population substantially.

This does not mean researchers must protect against every theoretically imaginable identification scenario. Under current UK anonymisation guidance, for example, the assessment considers means reasonably likely to be used, taking into account factors such as available information, technology, time, and cost. Other jurisdictions use their own legal standards.

The broader research lesson is that identifiability should be evaluated in the actual disclosure environment rather than imagined in a vacuum.

Pseudonymisation reduces risk but does not necessarily remove identifiability

Replacing names with participant codes can be an important safeguard. It separates obvious identifying information from the working dataset and can reduce unnecessary exposure.

But if additional information exists that reconnects the code to the person, the data have not necessarily become anonymous. Likewise, even without the code key, indirect information in the dataset may still permit identification.

Researchers should therefore distinguish between reducing identification risk and eliminating it. That distinction affects how identifiable research materials should be collected, accessed, shared, retained, and disposed of.

04 · A Practical Example

How ordinary demographic details can point to one participant

Hypothetical Example

A de-identified interview still contains a recognizable profile

A researcher removes the participant's name from an interview transcript. The transcript says that the participant is 37 years old, works as the principal of a rural secondary school in a named municipality, completed a doctorate overseas two years earlier, and recently received a nationally publicized education award.

Name removed The most obvious direct identifier is gone.
Individual details Age, occupation, municipality, education, and award history may each describe more than one person.
Combination Together, the details may narrow the description to one readily discoverable individual.
Interpretation The absence of a name does not establish that the transcript is anonymous.
Action The researcher considers whether some details can be generalized, suppressed, or otherwise managed without undermining the analysis and applies access controls if the material cannot appropriately be anonymised for the intended disclosure.

The example illustrates why identifiability is relational. What matters is not merely which fields remain but what those fields reveal together and what other information a likely recipient could connect to them.

05 · What Researchers Often Get Wrong

Common mistakes when deciding whether research data are identifiable

Misconception

Only names and contact details count as identifiers

Direct identifiers are the obvious cases, but location, online identifiers, physical characteristics, social information, occupational details, and many other factors may also permit direct or indirect identification.

Misconception

If no single variable identifies someone, the dataset is anonymous

Several weak identifiers can become powerful when combined. Assess combinations of variables and narrative details rather than testing each field in isolation.

Misconception

A photograph is identifiable only if the participant's name accompanies it

A recognizable face or other distinctive visual information may identify the person without any accompanying name. Whether participant photographs can be used in publications or presentations also raises separate consent and dissemination questions.

Misconception

Indirect identifiers have a fixed level of risk

The same detail can be uninformative in a large population and highly identifying in a small one. Occupation, age, location, and institutional affiliation become more or less revealing depending on context.

Misconception

If the general public cannot identify the participant, the data are anonymous

People within a participant's workplace, family, community, or specialist field may possess information that outsiders do not. The likely audience and available external information matter when assessing disclosure risk.

Misconception

Removing all context is always the safest solution

It may reduce disclosure risk, but it can also destroy the analytical value of qualitative or contextual data. Researchers need to balance data utility with confidentiality and consider restricted access or other safeguards when meaningful anonymisation would make the data unusable.

06 · What This Means for You

Ask who could connect the remaining clues to a real person

Instead of working from a rigid list of identifiers, examine how the information functions in your particular study. Start with obvious identifiers, then look for characteristics that distinguish participants and combinations that narrow the possible population.

A contextual identifiability test

If the information directly points to a particular person
Treat it as identifiable and apply the safeguards required for your study and jurisdiction.
If one characteristic substantially distinguishes the participant
Consider whether people with relevant knowledge could recognize or single out the participant.
If individual details appear harmless
Test their combinations and consider whether they can be linked with other reasonably available information.
If meaningful anonymisation would seriously damage the research value
Consider whether stronger access restrictions or another controlled disclosure arrangement is more appropriate than pretending the data are anonymous.
If you remain uncertain whether participants are identifiable
Use a cautious classification and consult the applicable institutional, ethics, data-protection, or legal guidance before broader disclosure.

Identifiability is therefore not merely a data-cleaning question. It is an assessment of the information, the population, the audience, the surrounding information environment, and the way the material will be released or accessed.

07 · A Quick Checklist

Before deciding that research data are non-identifiable

Check whether participants could be distinguished through:
Names, contact details, identification numbers, account names, or other direct identifiers.
Recognizable faces, voices, distinctive physical features, or other observable characteristics.
Specific locations, workplaces, schools, institutions, or geographic patterns.
Rare occupations, senior roles, unusual qualifications, or distinctive career histories.
Unique events, relationships, timelines, experiences, or narrative details.
Online identifiers, usernames, device information, or other persistent digital traces where relevant.
Combinations of details that are substantially more identifying together than separately.
External information reasonably available to the people who will access or receive the data.
A code key or other linkage mechanism that can reconnect pseudonymised records to participants.
08 · Frequently Asked Questions

Questions about indirect identifiers and research data

Can a person's face make research data identifiable?

Yes. A recognizable face can allow someone to identify the participant without a name. A photograph does not need to meet a narrower legal definition of biometric data to create an identification and confidentiality risk.

Can someone's voice identify them?

Yes. People familiar with a participant may recognize their voice, and the recording may contain additional contextual clues. Whether the recording is identifiable depends on the material and its disclosure context.

Is a city or region always identifying information?

No. Identifiability depends on context and granularity. A broad region may reveal little in one dataset, while a small locality combined with occupation, age, or another characteristic may substantially narrow who the participant could be.

Can age and occupation identify someone?

They can, particularly when combined with other information or when the relevant population is small. The appropriate question is not whether either variable is universally identifying but whether the available combination can distinguish the participant in context.

Does replacing names with participant codes make the dataset anonymous?

Not necessarily. If a key can reconnect codes to participants, the data remain linkable. Other information in the dataset may also permit identification independently of the key.

Can publicly available information make research data identifiable?

Yes. Information that appears non-identifying on its own may sometimes be linked with public or otherwise reasonably available sources. This is why identifiability assessments should consider realistic external information rather than only the research dataset in isolation.

Does anonymisation require zero theoretical possibility of identification?

Standards vary by jurisdiction. For example, current UK guidance evaluates whether identification is reasonably likely, considering objective factors such as available technology, time, and cost, rather than every purely hypothetical possibility. Researchers should use the standard applicable to their institution, study, and jurisdiction.

09 · The Bottom Line

What identifies someone depends on what the information reveals in context

The Bottom Line

Voice, face, location, occupation, institutional affiliation, digital identifiers, distinctive experiences, and many other characteristics can make research data identifiable, especially when several details can be combined or linked with outside information.

Do not equate removing names with anonymisation. Assess what remains, who could access it, what they could reasonably know or obtain, and whether the information can distinguish or reconnect the data to a particular participant.

10 · Sources and Further Reading

Authoritative sources on identifiers and research-data identifiability

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes