01 · The Question
Can the Same Dataset Be Anonymous to One Person but Identifiable to Another?
Imagine receiving a research dataset containing participant codes such as R104, R218, and R391. You have no names, contact information, or code key. Nothing available to you reveals who those participants are.
Elsewhere, however, another organisation holds the file connecting R104 to a particular person.
Are the data anonymous because you cannot identify anyone, or identifiable because someone else can?
In some legal and practical frameworks, the answer depends on whose hands the information is in and what additional information is realistically available to that party. But there are important limits to this idea.
03 · What You Need to Know
The Same Information Can Present Different Identification Possibilities to Different Parties
Identifiability Depends Partly on Available Information
Whether someone can identify a participant depends not only on the research file but also on what else that person can access.
A coded record such as P074 means little without context. Give someone a separate table mapping P074 to a name, and the situation changes immediately. Another recipient may not possess the table and may have no realistic way to obtain it.
This illustrates a broader principle: identifiability can depend on the relationship between the dataset, the recipient, and additional information.
Current ICO anonymisation guidance explicitly asks whether someone else, including an organisation receiving information, could identify people either from the information itself or by using other information it possesses or may obtain. The ICO refers to this as the "whose hands?" question.
The Data Holder and the Recipient May Not Be in the Same Position
Consider a university that holds a master research dataset containing participant identities. Before sharing data with an independent external research group, the university removes direct identifiers, transforms potentially identifying variables, withholds the linkage key, and imposes controls on the release.
The university still possesses information that connects the research records to participants. The external group may not.
It can therefore be misleading to describe the data simply as "anonymous" without specifying to whom and under what conditions. The source organisation and independent recipient may occupy materially different positions.
Source organisation
May possess identifiable source records, a code key, or other information that permits attribution.
Independent recipient
May receive a transformed dataset without the additional information needed to identify participants.
This Does Not Mean Every Coded Dataset Is Anonymous to Whoever Lacks the Key
Withholding the key is only one part of the analysis.
Suppose an external researcher receives coded data but the dataset also contains exact age, detailed occupation, precise location, and distinctive event dates. The researcher may be able to identify participants using public information without ever seeing the code key.
The recipient's inability to access one identification route does not prove that no other realistic route exists.
Researchers must still assess whether people can be identified without explicit names or a code key .
Additional Information Can Change the Status of the Same File
Imagine three people receiving exactly the same coded dataset.
Recipient
Additional Information
Identification Position
Research analyst
No key and no realistic external linkage source
May be unable to identify participants
Principal investigator
Has authorised access to the code key
Can reconnect coded records to participants
Participant's employer
Has detailed personnel records matching variables in the dataset
May be able to identify records through linkage even without the code key
The bytes in the research file have not changed. What changed is the information environment surrounding each recipient.
This is why re-identification risk can increase when other datasets are available .
Pseudonymized Data Illustrate the Problem Clearly
Pseudonymization intentionally separates identifying information from other data. Under GDPR-style frameworks, pseudonymized data remain personal data when they can be attributed to individuals through additional information.
The organisation holding the key therefore cannot simply call its coded dataset anonymous because analysts working with one copy cannot see participants' names.
UKRI's updated guidance on identifiability, anonymisation, and pseudonymisation is aimed specifically at research governance and emphasizes assessing the information that can make people identifiable and the challenges involved in keeping person-level research data anonymous.
The distinction between pseudonymized and anonymous information remains important precisely because a controlled identity route may still exist.
The "Whose Hands?" Approach Has Legal Limits
The idea that identifiability can differ by recipient should not be stretched into a universal shortcut.
The ICO states that its "whose hands?" approach applies in particular circumstances involving disclosure to an organisation that is not acting jointly with the disclosing organisation as a joint controller and is not its processor. It further explains that if information is personal data in a controller's hands, it remains personal data in the hands of that controller's processors, and the same status applies across joint controllers.
Watch Out
Do not conclude that a dataset becomes anonymous simply because one team member lacks access to the identity key. Researchers working for the same controller, processors acting on its behalf, joint controllers, and genuinely independent recipients may be treated differently under applicable data-protection law.
Research Regulations May Ask a Related but Different Question
Legal frameworks outside data-protection law may approach identifiability differently.
For example, US Office for Human Research Protections guidance concerning coded private information and biospecimens focuses in part on whether investigators can readily ascertain the identities of the individuals to whom coded information pertains. In certain circumstances, investigators receiving coded information without access to the key may not be considered to be working with individually identifiable information for relevant Common Rule purposes.
That does not mean the underlying source organisation has anonymous data, nor does it establish a universal definition for every research context. It shows why researchers must identify which regulatory question they are actually answering.
Contractual Restrictions Can Affect What a Recipient Can Realistically Do
A recipient's ability to identify participants is not determined solely by technical capability. Legal and organisational controls can also affect the realistic identification environment.
A data-sharing arrangement might prohibit attempts to re-identify participants, prohibit linkage with specified external datasets, restrict onward disclosure, limit access to named researchers, and require use within a secure environment.
Such controls can form part of risk management. They do not necessarily transform personal data into anonymous information under every legal framework, but they may be relevant when assessing realistic means of identification and the conditions under which information is disclosed.
Controlled Access Can Be More Appropriate Than Pretending Everyone Receives Anonymous Data
Some useful research datasets cannot be transformed sufficiently for unrestricted public release without destroying important analytical information.
In those cases, controlled access can preserve research utility while managing identification risk. NIST's guidance recognizes several sharing models, including public release, synthetic data, query interfaces, and protected environments, and recommends selecting the release model as part of the de-identification strategy.
This is often more intellectually honest than forcing all useful person-level data into a binary choice between "identified" and "publicly anonymous."
The Researcher Should Ask "Anonymous to Whom, and Under What Conditions?"
The word "anonymous" can conceal important differences unless its perspective is clear.
Before using it, ask:
Who holds the original identifiable information?
Who holds any linkage key?
Who receives the transformed dataset?
What other information does each party possess?
Can the recipient obtain additional information?
Are the parties independent, joint controllers, processors, or operating under another legal relationship?
What contractual and technical restrictions apply?
Could participants still be identified from the released data themselves?
Those questions provide a much more informative description of the research arrangement than simply declaring that "the dataset is anonymous."
06 · What This Means for You
Assess Identifiability Separately for Each Relevant Data Holder
When research information moves between organisations, do not ask only whether identifiers were removed. Map who possesses what.
A simple decision framework
If your organisation retains a code key or identifiable source records
Do not call its own coded data anonymous merely because routine analysts cannot access the key.
If an independent recipient receives transformed data without the key
Assess whether that recipient can realistically identify participants using the released information and other information available to it.
If the recipient has overlapping administrative or external data
Evaluate linkage and re-identification risk before treating the information as anonymous in that recipient's hands.
If the recipient is your processor or joint controller
Do not assume the "whose hands?" approach changes the information's legal status; verify the applicable data-protection rules.
If unrestricted sharing would create unacceptable identification risk
This approach is especially useful in collaborative research, repositories, secondary-data analysis, and multi-institutional projects. Rather than asking whether the file has one permanent label, document what each party receives, what additional information it controls, and what it is permitted and realistically able to do.
07 · A Quick Checklist
Ask Whose Hands the Research Data Are In
Before describing shared data as anonymous or non-identifiable, check:
Identify every organisation or research team that holds a copy of the data.
Document who holds the original identifiable records and any code or linkage key.
Determine what additional administrative, research, public, or commercial information each recipient possesses or can reasonably obtain.
Assess whether the released dataset itself contains indirect identifiers or distinctive combinations that permit identification without a key.
Clarify whether recipients are independent organisations, processors, joint controllers, collaborators, or operating under another relevant legal relationship.
Review contractual restrictions on re-identification, linkage, onward disclosure, and permitted uses.
Match the data-sharing model to the identification risk rather than assuming every useful dataset must be released publicly.
Verify the definitions required by the applicable data-protection law, human-subjects framework, ethics body, institution, and data-sharing agreement.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation