03 · What You Need to Know
Anonymity Depends on Whether People Remain Identifiable
Anonymous Does Not Mean "Contains No Names"
The central issue in anonymization is identifiability. The Information Commissioner's Office explains that a person can be identifiable without their name being known and that identifiability should be considered broadly. Relevant indicators include whether a person can be singled out and whether information can be linked.
This distinction matters because datasets often contain information that is not obviously identifying when viewed one variable at a time. Age, occupation, location, dates, household characteristics, institutional role, rare diagnoses, and other attributes may become identifying in combination.
A researcher therefore needs to consider how someone could be identified without a name. Simply stripping a conventional list of direct identifiers may reduce risk substantially, but it does not establish anonymity by itself.
Anonymity Is Contextual
The same dataset may present different identification risks in different hands or environments. A recipient who has no relevant additional information may be unable to connect records to individuals, while another person with access to a membership list, administrative database, social-media information, local knowledge, or another dataset may be able to do so.
The ICO's current anonymisation guidance therefore treats identifiability as context-specific and emphasizes considering both the information and who may gain access to it. The relevant question is not merely, "Can I identify these people?" but also whether other people who may obtain the information have means reasonably likely to enable identification.
This creates an important distinction explored further in whether data can be anonymous to one party but identifiable to another.
Absolute Zero Risk Is Not the Usual Test
If anonymity required researchers to prove that no person could ever be identified by any imaginable future method, almost no useful person-level dataset could confidently satisfy the standard.
Data-protection frameworks therefore commonly take a risk-based approach. Under the UK framework, the ICO states that researchers need not account for a purely hypothetical or theoretical possibility of identification. The assessment instead considers means reasonably likely to be used. Relevant factors can include cost, time, available technology, technological developments, access to additional information, and the circumstances in which data are released.
Under the GDPR framework, Recital 26 similarly directs attention to means reasonably likely to be used to identify a person and identifies objective considerations such as cost, time, available technology, and technological developments.
Watch Out
"Low risk of identification" and "anonymous" should not automatically be treated as synonyms. Whether a sufficiently low identification risk qualifies information as anonymous depends on the applicable legal, institutional, ethical, and technical standard. Document the basis for the classification rather than treating anonymity as a casual description.
Direct Identifiers Are Only the Beginning
A name, personal email address, telephone number, government identification number, or similar attribute can make identification straightforward. Removing such information is often an obvious first step.
The harder problem is indirect identification. Imagine a dataset describing a participant as a 67-year-old female dean at a small university in a particular province. No name appears, yet people familiar with the institution may immediately know who the record describes.
The risk depends on context. "Professor" may describe hundreds of people in one dataset. "The university's only professor of a rare specialty" may describe one. Identifiability is therefore not simply a property of individual columns.
Linkage Can Turn Ordinary Variables Into Identifiers
Researchers should also consider what happens when information is combined. A dataset containing broad demographic information may appear difficult to identify on its own. Another dataset may contain some of the same variables plus names. Matching overlapping characteristics can potentially reveal identities.
This is the logic behind linkage and re-identification risk. Additional information does not have to be contained in the research file itself. It may come from administrative records, public databases, publications, social media, professional directories, data breaches, or another research dataset.
The availability of external information is one reason re-identification risk can increase when datasets are combined.
Pseudonymized Data Are Not the Same as Anonymous Data
Suppose a researcher replaces every participant name with a randomly generated code and stores the name-to-code key in a separate encrypted file. This is a useful protection because someone who obtains the research dataset alone no longer sees the participants' names.
But the identity connection still exists. A person with legitimate or illegitimate access to both the dataset and the key could restore it.
Under the GDPR and UK GDPR frameworks, data that can be attributed to a person through separately held additional information are treated as pseudonymized rather than anonymous. Pseudonymization is an important safeguard, but it does not remove the data from the personal-data regime merely because the linking information is stored elsewhere.
Pseudonymized data
Direct identifiers may be replaced or separated, but additional information still permits attribution to individuals.
Anonymous information
The information does not relate to an identified or identifiable person under the applicable standard and context.
Different Types of Data Create Different Anonymization Problems
Anonymizing a spreadsheet is not the same problem as anonymizing an interview transcript, photograph, genome, or audio recording.
Structured quantitative data may permit generalisation, suppression, aggregation, randomisation, or other statistical disclosure-control techniques. The ICO, for example, discusses generalisation and randomisation as approaches to reducing identifiability and notes that masking alone is not necessarily sufficient as an anonymisation technique.
Qualitative data pose different challenges. Removing a participant's name from an interview transcript may leave descriptions of workplaces, relationships, events, personal histories, quotations, and other contextual clues. Images and video may contain faces, locations, badges, signs, or distinctive environments. Audio can contain recognisable voices as well as identifying speech.
Anonymization therefore needs to address the information actually contained in the research material, not merely the file format or a predetermined checklist of fields.
Public Release Raises a Different Risk Than Controlled Access
Who receives the data matters. Information released openly can potentially be examined by an unknown number of people, combined with diverse external sources, copied indefinitely, and subjected to new identification techniques.
A controlled research environment can restrict who accesses information, what additional datasets they can use, what analyses they can perform, and what outputs leave the environment. Those controls can materially affect identification risk.
The ICO consequently recommends considering the release model when assessing anonymization, distinguishing circumstances such as internal use, controlled disclosure to another organisation or defined group, and release to the wider public.
Researchers planning open-data sharing should therefore not assume that a dataset suitable for controlled access is automatically suitable for unrestricted public release.
Anonymity Can Change Over Time
An anonymization assessment is not necessarily permanent. New datasets may become publicly available. Computational techniques may improve. Information once difficult or expensive to obtain may become readily accessible. A recipient may gain access to new linkage sources.
The ICO specifically recommends reassessing identification risk when new vulnerabilities, new datasets, new recipients, changed purposes, or technological developments alter the circumstances.
This does not mean anonymous data inevitably become identifiable. It means anonymity is an evidence-based judgment made within a particular information environment and should be reconsidered when that environment materially changes.
06 · What This Means for You
Treat Anonymity as a Defensible Assessment, Not a Checkbox
Before describing research data as anonymous, work through the ways a person could realistically be identified. Begin with the data themselves, then widen the assessment to include additional information, likely recipients, the release environment, and plausible identification methods.
A simple decision framework
If a direct identifier remains
The relevant record is generally not anonymous. Determine whether that identifier is necessary and how it should be protected.
If direct identifiers are removed
Assess indirect identifiers, singling out, linkability, contextual clues, and realistic external information before concluding that anonymity has been achieved.
If a key or other additional information can restore identities
Treat the distinction between pseudonymization and anonymization carefully and apply the appropriate safeguards.
If data will be publicly released
Assess identification risk for an open environment rather than assuming that protections adequate for a controlled research team remain adequate after public disclosure.
If the information environment materially changes
Reassess identifiability rather than relying indefinitely on the original anonymization decision.
Where the remaining identification risk cannot be reduced sufficiently without destroying the research value of the data, the answer need not be to pretend the data are anonymous. Controlled access, pseudonymization, separation of identifiers, access restrictions, data-use agreements, secure environments, and other safeguards may allow useful research while acknowledging that the information remains identifiable.
Sometimes the most rigorous statement a researcher can make is not "these data are anonymous" but "these data remain potentially identifiable, and here is how that risk is managed." Less glamorous, perhaps, but methods sections have survived worse.
07 · A Quick Checklist
Before Calling Research Data Anonymous, Check the Identification Risk
Before classifying data as anonymous, check:
Identify and remove or appropriately transform unnecessary direct identifiers.
Review dates, locations, occupations, institutional roles, rare characteristics, free text, images, audio, and metadata for identifying information.
Determine whether a code key, lookup table, contact file, or other additional information can reconnect records to participants.
Consider what external information likely recipients could reasonably use for linkage or identification.
Assess whether individuals can be singled out even if their real-world names are unknown.
Evaluate the actual release environment, distinguishing controlled access from unrestricted public release.
Document the anonymization techniques, assumptions, remaining risks, and basis for concluding that identification is sufficiently remote under the applicable standard.
Reassess identification risk when new datasets, technologies, recipients, or uses materially change the context.