Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Research Data Be Anonymous to the Researcher but Identifiable to Someone Else?

Whether research data are identifiable can sometimes depend on whose hands they are in. A researcher may lack any realistic way to identify participants while another organisation can identify the same records using a key or additional information.

279
Anonymous to One Person, Identifiable to Another Guide 279 of 398
01 · The Question

Can the Same Dataset Be Anonymous to One Person but Identifiable to Another?

Imagine receiving a research dataset containing participant codes such as R104, R218, and R391. You have no names, contact information, or code key. Nothing available to you reveals who those participants are.

Elsewhere, however, another organisation holds the file connecting R104 to a particular person.

Are the data anonymous because you cannot identify anyone, or identifiable because someone else can?

In some legal and practical frameworks, the answer depends on whose hands the information is in and what additional information is realistically available to that party. But there are important limits to this idea.

02 · The Short Answer

Identifiability Can Depend on Who Holds the Data

In Brief

Yes, in some circumstances information may be effectively anonymous or non-identifiable to one independent recipient while remaining identifiable or personal data in another party's hands because that other party possesses additional information or means of identification.

This is sometimes called the "whose hands?" question. It is not a universal rule that automatically applies to every researcher, processor, collaborator, or jurisdiction, so the relevant legal, ethical, and institutional framework must be checked before classifying the data.

03 · What You Need to Know

The Same Information Can Present Different Identification Possibilities to Different Parties

Identifiability Depends Partly on Available Information

Whether someone can identify a participant depends not only on the research file but also on what else that person can access.

A coded record such as P074 means little without context. Give someone a separate table mapping P074 to a name, and the situation changes immediately. Another recipient may not possess the table and may have no realistic way to obtain it.

This illustrates a broader principle: identifiability can depend on the relationship between the dataset, the recipient, and additional information.

Current ICO anonymisation guidance explicitly asks whether someone else, including an organisation receiving information, could identify people either from the information itself or by using other information it possesses or may obtain. The ICO refers to this as the "whose hands?" question.

The Data Holder and the Recipient May Not Be in the Same Position

Consider a university that holds a master research dataset containing participant identities. Before sharing data with an independent external research group, the university removes direct identifiers, transforms potentially identifying variables, withholds the linkage key, and imposes controls on the release.

The university still possesses information that connects the research records to participants. The external group may not.

It can therefore be misleading to describe the data simply as "anonymous" without specifying to whom and under what conditions. The source organisation and independent recipient may occupy materially different positions.

Source organisation May possess identifiable source records, a code key, or other information that permits attribution.
Independent recipient May receive a transformed dataset without the additional information needed to identify participants.

This Does Not Mean Every Coded Dataset Is Anonymous to Whoever Lacks the Key

Withholding the key is only one part of the analysis.

Suppose an external researcher receives coded data but the dataset also contains exact age, detailed occupation, precise location, and distinctive event dates. The researcher may be able to identify participants using public information without ever seeing the code key.

The recipient's inability to access one identification route does not prove that no other realistic route exists.

Researchers must still assess whether people can be identified without explicit names or a code key.

Additional Information Can Change the Status of the Same File

Imagine three people receiving exactly the same coded dataset.

Recipient Additional Information Identification Position
Research analyst No key and no realistic external linkage source May be unable to identify participants
Principal investigator Has authorised access to the code key Can reconnect coded records to participants
Participant's employer Has detailed personnel records matching variables in the dataset May be able to identify records through linkage even without the code key

The bytes in the research file have not changed. What changed is the information environment surrounding each recipient.

This is why re-identification risk can increase when other datasets are available.

Pseudonymized Data Illustrate the Problem Clearly

Pseudonymization intentionally separates identifying information from other data. Under GDPR-style frameworks, pseudonymized data remain personal data when they can be attributed to individuals through additional information.

The organisation holding the key therefore cannot simply call its coded dataset anonymous because analysts working with one copy cannot see participants' names.

UKRI's updated guidance on identifiability, anonymisation, and pseudonymisation is aimed specifically at research governance and emphasizes assessing the information that can make people identifiable and the challenges involved in keeping person-level research data anonymous.

The distinction between pseudonymized and anonymous information remains important precisely because a controlled identity route may still exist.

The "Whose Hands?" Approach Has Legal Limits

The idea that identifiability can differ by recipient should not be stretched into a universal shortcut.

The ICO states that its "whose hands?" approach applies in particular circumstances involving disclosure to an organisation that is not acting jointly with the disclosing organisation as a joint controller and is not its processor. It further explains that if information is personal data in a controller's hands, it remains personal data in the hands of that controller's processors, and the same status applies across joint controllers.

Watch Out

Do not conclude that a dataset becomes anonymous simply because one team member lacks access to the identity key. Researchers working for the same controller, processors acting on its behalf, joint controllers, and genuinely independent recipients may be treated differently under applicable data-protection law.

Research Regulations May Ask a Related but Different Question

Legal frameworks outside data-protection law may approach identifiability differently.

For example, US Office for Human Research Protections guidance concerning coded private information and biospecimens focuses in part on whether investigators can readily ascertain the identities of the individuals to whom coded information pertains. In certain circumstances, investigators receiving coded information without access to the key may not be considered to be working with individually identifiable information for relevant Common Rule purposes.

That does not mean the underlying source organisation has anonymous data, nor does it establish a universal definition for every research context. It shows why researchers must identify which regulatory question they are actually answering.

Contractual Restrictions Can Affect What a Recipient Can Realistically Do

A recipient's ability to identify participants is not determined solely by technical capability. Legal and organisational controls can also affect the realistic identification environment.

A data-sharing arrangement might prohibit attempts to re-identify participants, prohibit linkage with specified external datasets, restrict onward disclosure, limit access to named researchers, and require use within a secure environment.

Such controls can form part of risk management. They do not necessarily transform personal data into anonymous information under every legal framework, but they may be relevant when assessing realistic means of identification and the conditions under which information is disclosed.

Controlled Access Can Be More Appropriate Than Pretending Everyone Receives Anonymous Data

Some useful research datasets cannot be transformed sufficiently for unrestricted public release without destroying important analytical information.

In those cases, controlled access can preserve research utility while managing identification risk. NIST's guidance recognizes several sharing models, including public release, synthetic data, query interfaces, and protected environments, and recommends selecting the release model as part of the de-identification strategy.

This is often more intellectually honest than forcing all useful person-level data into a binary choice between "identified" and "publicly anonymous."

The Researcher Should Ask "Anonymous to Whom, and Under What Conditions?"

The word "anonymous" can conceal important differences unless its perspective is clear.

Before using it, ask:

  • Who holds the original identifiable information?
  • Who holds any linkage key?
  • Who receives the transformed dataset?
  • What other information does each party possess?
  • Can the recipient obtain additional information?
  • Are the parties independent, joint controllers, processors, or operating under another legal relationship?
  • What contractual and technical restrictions apply?
  • Could participants still be identified from the released data themselves?

Those questions provide a much more informative description of the research arrangement than simply declaring that "the dataset is anonymous."

04 · A Practical Example

One Dataset, Three Different Identification Environments

Hypothetical Example

A Multi-University Student Well-Being Project

University A collects identifiable student information for a longitudinal study. Each participant receives a random research code. The university stores the code key separately.

University A Authorised staff can access the code key and reconnect research records to individual students. The coded records are therefore not anonymous to the university merely because the analysis file contains no names.
Independent research centre The university prepares a transformed dataset for an independent centre. The centre receives no key, no direct identifiers, less precise demographics, and no contractual or practical route to obtain the source records.
The centre's position Depending on the governing legal framework and the residual information, the dataset may be effectively anonymous or non-identifiable in the centre's hands even though University A retains identifiable source information.
A different recipient Suppose the same file were instead sent to an organisation possessing detailed student administrative records containing the same demographic and enrolment variables.
The risk changes That organisation may have a realistic linkage route unavailable to the independent centre. The same released variables now exist in a different identification environment.

The example illustrates why identifiability should be assessed in relation to both the information and the relevant holder. It also shows why data-sharing decisions should be based on the actual recipient rather than a generic label applied once to a file.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Relative Identifiability

Misconception

If I Cannot Identify Participants, the Data Are Anonymous Everywhere

Another party may hold a key, source records, administrative information, or other data that permit identification. Your inability to identify participants does not automatically establish the status of the information in everyone else's hands.

Misconception

If the Principal Investigator Has the Key, the Dataset Is Anonymous to the Rest of the Team

That conclusion may be incorrect, particularly when researchers operate under the same controller or organisational arrangement. Restricted access is valuable, but internal access controls do not automatically change the legal status of the data for each individual staff member.

Misconception

No Code Key Means No Identification Risk

The recipient may still identify participants through indirect identifiers, contextual knowledge, public information, or other datasets. Withholding the key removes one route, not necessarily every route.

Misconception

The Same Dataset Must Have Exactly the Same Status for Every Organisation

Some frameworks allow the identifiability analysis to depend on what an independent recipient can realistically identify, although the legal relationship among parties matters. Do not generalise this principle beyond the framework being applied.

Misconception

A Data-Sharing Agreement Automatically Makes Data Anonymous

Contractual restrictions can reduce realistic identification risk and prohibit re-identification, but a contract does not magically erase identifying information. Technical characteristics, additional information, legal relationships, and the governing standard still matter.

06 · What This Means for You

Assess Identifiability Separately for Each Relevant Data Holder

When research information moves between organisations, do not ask only whether identifiers were removed. Map who possesses what.

A simple decision framework

If your organisation retains a code key or identifiable source records
Do not call its own coded data anonymous merely because routine analysts cannot access the key.
If an independent recipient receives transformed data without the key
Assess whether that recipient can realistically identify participants using the released information and other information available to it.
If the recipient has overlapping administrative or external data
Evaluate linkage and re-identification risk before treating the information as anonymous in that recipient's hands.
If the recipient is your processor or joint controller
Do not assume the "whose hands?" approach changes the information's legal status; verify the applicable data-protection rules.
If unrestricted sharing would create unacceptable identification risk
Consider a controlled-access model and restrict who can access identifiable or potentially identifiable information.

This approach is especially useful in collaborative research, repositories, secondary-data analysis, and multi-institutional projects. Rather than asking whether the file has one permanent label, document what each party receives, what additional information it controls, and what it is permitted and realistically able to do.

07 · A Quick Checklist

Ask Whose Hands the Research Data Are In

Before describing shared data as anonymous or non-identifiable, check:
Identify every organisation or research team that holds a copy of the data.
Document who holds the original identifiable records and any code or linkage key.
Determine what additional administrative, research, public, or commercial information each recipient possesses or can reasonably obtain.
Assess whether the released dataset itself contains indirect identifiers or distinctive combinations that permit identification without a key.
Clarify whether recipients are independent organisations, processors, joint controllers, collaborators, or operating under another relevant legal relationship.
Review contractual restrictions on re-identification, linkage, onward disclosure, and permitted uses.
Match the data-sharing model to the identification risk rather than assuming every useful dataset must be released publicly.
Verify the definitions required by the applicable data-protection law, human-subjects framework, ethics body, institution, and data-sharing agreement.
08 · Frequently Asked Questions

Frequently Asked Questions About Anonymous Data in Different Hands

Can data be anonymous to a secondary researcher if the original researcher knows participants' identities?

Potentially, depending on the governing framework, the relationship between the parties, what the secondary researcher receives, and whether that researcher has realistic means of identifying participants. The source researcher's possession of identifiable records should still be described accurately.

Are coded data anonymous to an analyst who cannot access the key?

Not automatically. The analyst may still be operating within an organisation that holds the key, or the data themselves may permit identification through other information. Restricted key access is a safeguard, not a universal declaration of anonymity.

Can an external researcher receive anonymous data while the university retains identifiable data?

In some circumstances, yes. An independent recipient may receive sufficiently transformed information without realistic access to the additional information needed for identification. Whether it legally qualifies as anonymous must be assessed under the applicable framework.

Does a data-use agreement make data anonymous?

No. It may restrict re-identification, linkage, disclosure, and other activities and thereby contribute to risk management. The actual information, available additional data, recipient relationship, and governing legal standard still determine its status.

Does withholding the participant code key make data anonymous?

Not by itself. The remaining variables may still identify participants directly or through linkage with other information. The recipient's access to the key is only one part of the identifiability assessment.

Can a processor treat data as anonymous because it cannot identify participants?

Under the ICO's UK GDPR guidance, the "whose hands?" approach does not operate that way for processors. A processor processes personal data on the controller's behalf, so the status of the information in the controller's hands remains relevant. Other jurisdictions should be checked separately.

Why does this matter for research repositories?

Repositories may serve different audiences and access models. Data unsuitable for unrestricted public release may sometimes be made available through controlled access, with restrictions on recipients, linkage, use, and outputs. Identifiability should be assessed for the actual sharing arrangement.

09 · The Bottom Line

Identifiability Can Depend on Whose Hands Hold the Data

The Bottom Line

The same research information may sometimes be non-identifiable to one independent recipient while remaining identifiable to another party that possesses a code key, source records, contextual knowledge, or other information enabling identification.

Do not turn this into a shortcut for calling coded data anonymous. Ask who holds the information, what else each party can access, whether the recipient is genuinely independent, and which legal or ethical definition applies. In data governance, "who knows what?" is often a more useful question than "what is this file called?"

10 · Sources and Further Reading

Authoritative Sources on Relative Identifiability

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes