Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Research Data Ever Be Truly Anonymous?

Research data can sometimes be effectively anonymous, but anonymity is not achieved simply by deleting names. Whether people remain identifiable depends on the data, available external information, who has access, and the realistic means of identification.

273
Can Research Data Be Truly Anonymous? Guide 273 of 398
01 · The Question

Can Researchers Really Make Data Anonymous?

Delete the names. Remove the email addresses. Replace exact ages with age ranges. Is the dataset now anonymous?

Perhaps. But anonymity is harder to establish than the absence of obvious identifiers. A person may still be identifiable from an unusual combination of characteristics, a distinctive quotation, location information, dates, linked public records, or other information available to whoever receives the data.

This creates an uncomfortable but important question: if identification can never be ruled out with metaphysical certainty, can researchers legitimately call any data anonymous?

02 · The Short Answer

Yes, but Anonymity Is Based on Realistic Identifiability, Not Absolute Impossibility

In Brief

Research data can be effectively anonymous when people are no longer identified or identifiable under the applicable standard, but anonymity generally does not require proving that identification is impossible under every imaginable circumstance.

Researchers must assess realistic identification risk in context. That includes the information in the dataset, other information that could reasonably be available, the people who may receive it, the techniques and resources that could be used, and how these factors may change over time.

03 · What You Need to Know

Anonymity Depends on Whether People Remain Identifiable

Anonymous Does Not Mean "Contains No Names"

The central issue in anonymization is identifiability. The Information Commissioner's Office explains that a person can be identifiable without their name being known and that identifiability should be considered broadly. Relevant indicators include whether a person can be singled out and whether information can be linked.

This distinction matters because datasets often contain information that is not obviously identifying when viewed one variable at a time. Age, occupation, location, dates, household characteristics, institutional role, rare diagnoses, and other attributes may become identifying in combination.

A researcher therefore needs to consider how someone could be identified without a name. Simply stripping a conventional list of direct identifiers may reduce risk substantially, but it does not establish anonymity by itself.

Anonymity Is Contextual

The same dataset may present different identification risks in different hands or environments. A recipient who has no relevant additional information may be unable to connect records to individuals, while another person with access to a membership list, administrative database, social-media information, local knowledge, or another dataset may be able to do so.

The ICO's current anonymisation guidance therefore treats identifiability as context-specific and emphasizes considering both the information and who may gain access to it. The relevant question is not merely, "Can I identify these people?" but also whether other people who may obtain the information have means reasonably likely to enable identification.

This creates an important distinction explored further in whether data can be anonymous to one party but identifiable to another.

Absolute Zero Risk Is Not the Usual Test

If anonymity required researchers to prove that no person could ever be identified by any imaginable future method, almost no useful person-level dataset could confidently satisfy the standard.

Data-protection frameworks therefore commonly take a risk-based approach. Under the UK framework, the ICO states that researchers need not account for a purely hypothetical or theoretical possibility of identification. The assessment instead considers means reasonably likely to be used. Relevant factors can include cost, time, available technology, technological developments, access to additional information, and the circumstances in which data are released.

Under the GDPR framework, Recital 26 similarly directs attention to means reasonably likely to be used to identify a person and identifies objective considerations such as cost, time, available technology, and technological developments.

Watch Out

"Low risk of identification" and "anonymous" should not automatically be treated as synonyms. Whether a sufficiently low identification risk qualifies information as anonymous depends on the applicable legal, institutional, ethical, and technical standard. Document the basis for the classification rather than treating anonymity as a casual description.

Direct Identifiers Are Only the Beginning

A name, personal email address, telephone number, government identification number, or similar attribute can make identification straightforward. Removing such information is often an obvious first step.

The harder problem is indirect identification. Imagine a dataset describing a participant as a 67-year-old female dean at a small university in a particular province. No name appears, yet people familiar with the institution may immediately know who the record describes.

The risk depends on context. "Professor" may describe hundreds of people in one dataset. "The university's only professor of a rare specialty" may describe one. Identifiability is therefore not simply a property of individual columns.

Linkage Can Turn Ordinary Variables Into Identifiers

Researchers should also consider what happens when information is combined. A dataset containing broad demographic information may appear difficult to identify on its own. Another dataset may contain some of the same variables plus names. Matching overlapping characteristics can potentially reveal identities.

This is the logic behind linkage and re-identification risk. Additional information does not have to be contained in the research file itself. It may come from administrative records, public databases, publications, social media, professional directories, data breaches, or another research dataset.

The availability of external information is one reason re-identification risk can increase when datasets are combined.

Pseudonymized Data Are Not the Same as Anonymous Data

Suppose a researcher replaces every participant name with a randomly generated code and stores the name-to-code key in a separate encrypted file. This is a useful protection because someone who obtains the research dataset alone no longer sees the participants' names.

But the identity connection still exists. A person with legitimate or illegitimate access to both the dataset and the key could restore it.

Under the GDPR and UK GDPR frameworks, data that can be attributed to a person through separately held additional information are treated as pseudonymized rather than anonymous. Pseudonymization is an important safeguard, but it does not remove the data from the personal-data regime merely because the linking information is stored elsewhere.

Pseudonymized data Direct identifiers may be replaced or separated, but additional information still permits attribution to individuals.
Anonymous information The information does not relate to an identified or identifiable person under the applicable standard and context.

Different Types of Data Create Different Anonymization Problems

Anonymizing a spreadsheet is not the same problem as anonymizing an interview transcript, photograph, genome, or audio recording.

Structured quantitative data may permit generalisation, suppression, aggregation, randomisation, or other statistical disclosure-control techniques. The ICO, for example, discusses generalisation and randomisation as approaches to reducing identifiability and notes that masking alone is not necessarily sufficient as an anonymisation technique.

Qualitative data pose different challenges. Removing a participant's name from an interview transcript may leave descriptions of workplaces, relationships, events, personal histories, quotations, and other contextual clues. Images and video may contain faces, locations, badges, signs, or distinctive environments. Audio can contain recognisable voices as well as identifying speech.

Anonymization therefore needs to address the information actually contained in the research material, not merely the file format or a predetermined checklist of fields.

Public Release Raises a Different Risk Than Controlled Access

Who receives the data matters. Information released openly can potentially be examined by an unknown number of people, combined with diverse external sources, copied indefinitely, and subjected to new identification techniques.

A controlled research environment can restrict who accesses information, what additional datasets they can use, what analyses they can perform, and what outputs leave the environment. Those controls can materially affect identification risk.

The ICO consequently recommends considering the release model when assessing anonymization, distinguishing circumstances such as internal use, controlled disclosure to another organisation or defined group, and release to the wider public.

Researchers planning open-data sharing should therefore not assume that a dataset suitable for controlled access is automatically suitable for unrestricted public release.

Anonymity Can Change Over Time

An anonymization assessment is not necessarily permanent. New datasets may become publicly available. Computational techniques may improve. Information once difficult or expensive to obtain may become readily accessible. A recipient may gain access to new linkage sources.

The ICO specifically recommends reassessing identification risk when new vulnerabilities, new datasets, new recipients, changed purposes, or technological developments alter the circumstances.

This does not mean anonymous data inevitably become identifiable. It means anonymity is an evidence-based judgment made within a particular information environment and should be reconsidered when that environment materially changes.

04 · A Practical Example

Why Removing Names May Still Leave Someone Identifiable

Hypothetical Example

An Interview Dataset From a Small Academic Department

A researcher interviews 18 employees about workplace culture. Before sharing transcripts with another research team, all names and email addresses are deleted.

Initial impression The transcripts contain no names. The researcher initially considers them anonymous.
Look at the remaining context One participant describes being the department's only full professor hired in a particular year, mentions a distinctive administrative role, and discusses a publicly reported project.
Consider outside information The university website identifies who holds that administrative role, while public professional profiles reveal appointment histories.
Assess identifiability A motivated person familiar with the institution may be able to connect the transcript to the participant even though no direct identifier appears.
Reduce the risk The researcher could consider removing unnecessary contextual details, generalising dates and roles, paraphrasing or withholding particularly identifying passages where methodologically and ethically appropriate, or using controlled access rather than public release.

The lesson is not that qualitative data can never be anonymized. It is that deleting names addresses only the most obvious identification route. Effective anonymization requires researchers to ask what a recipient could infer from everything that remains.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Truly Anonymous Data

Misconception

Anonymous Means Identification Must Be Mathematically Impossible

Not under every framework. Risk-based approaches such as the ICO's focus on whether identification is sufficiently remote and on means reasonably likely to be used, rather than every purely theoretical possibility. Researchers should nevertheless verify the standard applicable to their study rather than importing one jurisdiction's test universally.

Misconception

Removing Names Is Anonymization

It may be one step, but indirect identifiers, distinctive combinations, contextual details, metadata, and external information can still permit identification. Names are not the boundary between personal and anonymous information.

Misconception

A Code Makes the Dataset Anonymous

If a code can be linked back to a participant through additional information, the data may be pseudonymized rather than anonymous. Keeping the key in another file improves protection but does not necessarily eliminate identifiability.

Misconception

If the Researcher Cannot Identify Participants, Nobody Can

Another recipient may possess different knowledge or additional datasets. Anonymization assessments should consider relevant parties who may obtain the information and the realistic means available to them, not only the researcher's personal ability to recognise participants.

Misconception

An Anonymous Dataset Stays Anonymous Forever

Identification risk can change when technology, public information, available datasets, recipients, or uses change. Significant changes may warrant reassessment, particularly before wider data release or new linkage.

Misconception

More Anonymization Is Always Better

Anonymization can reduce privacy risk, but aggressive generalisation, suppression, perturbation, or removal of information may also reduce analytical usefulness. Researchers often need to manage a trade-off between data utility and identification risk while meeting applicable ethical and legal requirements.

06 · What This Means for You

Treat Anonymity as a Defensible Assessment, Not a Checkbox

Before describing research data as anonymous, work through the ways a person could realistically be identified. Begin with the data themselves, then widen the assessment to include additional information, likely recipients, the release environment, and plausible identification methods.

A simple decision framework

If a direct identifier remains
The relevant record is generally not anonymous. Determine whether that identifier is necessary and how it should be protected.
If direct identifiers are removed
Assess indirect identifiers, singling out, linkability, contextual clues, and realistic external information before concluding that anonymity has been achieved.
If a key or other additional information can restore identities
Treat the distinction between pseudonymization and anonymization carefully and apply the appropriate safeguards.
If data will be publicly released
Assess identification risk for an open environment rather than assuming that protections adequate for a controlled research team remain adequate after public disclosure.
If the information environment materially changes
Reassess identifiability rather than relying indefinitely on the original anonymization decision.

Where the remaining identification risk cannot be reduced sufficiently without destroying the research value of the data, the answer need not be to pretend the data are anonymous. Controlled access, pseudonymization, separation of identifiers, access restrictions, data-use agreements, secure environments, and other safeguards may allow useful research while acknowledging that the information remains identifiable.

Sometimes the most rigorous statement a researcher can make is not "these data are anonymous" but "these data remain potentially identifiable, and here is how that risk is managed." Less glamorous, perhaps, but methods sections have survived worse.

07 · A Quick Checklist

Before Calling Research Data Anonymous, Check the Identification Risk

Before classifying data as anonymous, check:
Identify and remove or appropriately transform unnecessary direct identifiers.
Examine demographic details that could identify participants individually or in combination.
Review dates, locations, occupations, institutional roles, rare characteristics, free text, images, audio, and metadata for identifying information.
Determine whether a code key, lookup table, contact file, or other additional information can reconnect records to participants.
Consider what external information likely recipients could reasonably use for linkage or identification.
Assess whether individuals can be singled out even if their real-world names are unknown.
Evaluate the actual release environment, distinguishing controlled access from unrestricted public release.
Document the anonymization techniques, assumptions, remaining risks, and basis for concluding that identification is sufficiently remote under the applicable standard.
Reassess identification risk when new datasets, technologies, recipients, or uses materially change the context.
08 · Frequently Asked Questions

Frequently Asked Questions About Anonymous Research Data

Does anonymous mean there is absolutely zero chance of identification?

Not necessarily. Some regulatory frameworks assess whether identification is realistically or reasonably likely rather than requiring proof against every theoretical possibility. The precise threshold depends on the applicable legal and institutional framework.

Are data anonymous once names and email addresses are removed?

Not automatically. Other variables, combinations of attributes, contextual information, metadata, or external datasets may still permit participants to be identified or singled out.

Can qualitative interview data ever be anonymous?

Potentially, but qualitative material can be difficult to anonymize because narratives may contain distinctive events, relationships, locations, occupations, quotations, and other contextual information. Researchers need to assess the entire informational content rather than merely redact names.

Are pseudonymized data anonymous?

No, not when additional information can still be used to attribute the data to individuals under the applicable framework. Pseudonymization can be a valuable safeguard while the information remains personal or identifiable.

Can a dataset be anonymous to me but identifiable to someone else?

Yes, depending on the legal framework and circumstances. Different parties may have different additional information, technical capabilities, or contextual knowledge. Identifiability should therefore be assessed in relation to the relevant recipients and data environment.

Does aggregation guarantee anonymity?

No. Aggregation can reduce identification risk, but very small groups, rare characteristics, differencing between tables, or other disclosure patterns may still reveal information about individuals. The level and structure of aggregation matter.

Can data become identifiable again after being considered anonymous?

Identification risk can change as new information, technologies, linkage opportunities, or recipients emerge. Where circumstances materially change, researchers should reconsider whether the original anonymization assessment remains defensible.

Should all research data be anonymized?

No. Some research requires participant linkage, follow-up, longitudinal analysis, or other uses incompatible with complete anonymization. Researchers should minimise unnecessary identification and use safeguards proportionate to the actual risks and requirements rather than assuming one data state suits every study.

09 · The Bottom Line

Truly Anonymous Does Not Mean Immune to Every Imaginary Identification Attempt

The Bottom Line

Research data can be effectively anonymous when individuals are no longer identifiable under the applicable standard, but deleting names or direct identifiers alone is not enough to establish that conclusion.

Assess the whole identification environment: the remaining variables, possible linkage, contextual clues, available external information, likely recipients, release model, technology, and how these may change. If meaningful identification risk remains, describe and protect the data accordingly rather than stretching the word "anonymous."

10 · Sources and Further Reading

Authoritative Sources on Anonymization and Identifiability

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes