Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Does De-Identification Automatically Make Secondary Use Ethically Acceptable?

De-identification can substantially reduce privacy risks and may change how secondary research is regulated, but it does not automatically make every secondary use ethically acceptable. Consent commitments, purpose, residual risks, governance, and potential harms can still matter.

370
Does De-Identification Make Secondary Use Ethical? Guide 370 of 398
01 · The Question

If Researchers Cannot Identify Participants, Is the Ethical Problem Solved?

Suppose you want to reuse an existing dataset, but the original consent does not clearly cover your new study. Someone proposes an apparently simple solution: remove the names, identification numbers, and other identifiers, then use the data as de-identified information.

De-identification can make an important ethical and regulatory difference. But does it transform a questionable secondary use into an acceptable one simply because the researcher can no longer readily identify the people represented in the dataset?

02 · The Short Answer

De-Identification Reduces Some Risks, but It Does Not Settle Every Ethical Question

In Brief

No. De-identification can substantially reduce privacy risk and may change the regulatory status of secondary research, but it does not automatically make every secondary use ethically acceptable.

You still need to consider whether the data are genuinely non-identifiable under the applicable framework, whether the proposed use conflicts with commitments made to participants, whether re-identification or group-level harms remain plausible, and whether repository, legal, contractual, or institutional restrictions continue to apply.

03 · What You Need to Know

Identifiability Is One Part of the Ethical Analysis, Not the Whole of It

De-identification can materially change the regulatory analysis

Identifiability matters because many human-subjects and privacy frameworks impose different requirements depending on whether researchers can identify the people behind the information.

Under the U.S. Common Rule, for example, a human subject includes a living individual about whom an investigator obtains, uses, studies, analyzes, or generates identifiable private information or identifiable biospecimens. OHRP states that secondary research involving only coded private information or coded biospecimens may not involve human subjects when investigators cannot readily ascertain the identities of the individuals because access to the code key is appropriately restricted.

The Common Rule also contains an exemption for certain secondary research in which information is recorded so that subjects' identities cannot readily be ascertained directly or through linked identifiers, provided that investigators do not contact or re-identify subjects.

De-identified, coded, and anonymized do not necessarily mean the same thing

Researchers often use these terms loosely, but regulatory definitions can differ. A coded dataset may have names replaced by numbers while a separate key still connects those numbers to individuals. Whether that dataset is considered identifiable can depend on who possesses the key, whether the research team can obtain it, and which regulatory framework applies.

Direct identifiers removed Names or obvious identifiers have been removed, but other variables or a code key may still permit identification.
Coded data Identifiers have been replaced by a code and a key exists somewhere that can reconnect the code to the individual.
Non-identifiable to the investigator The investigator cannot readily ascertain identities directly, through a coding system, or through other reasonably available means under the applicable standard.

Do not decide which category applies merely by looking at whether the spreadsheet contains a “Name” column. The actual data, access arrangements, available external information, and governing definitions matter.

Regulatory status and ethical acceptability are different questions

A study can fall outside a particular human-subjects regulation without becoming ethically neutral. OHRP's own advisory material illustrates the distinction clearly in the context of biospecimens: a secondary use of non-readily-identifiable material may fall outside Common Rule human-subjects requirements, while researchers and institutions may still have an obligation to honor agreements they made with participants about how their material would be used.

This distinction is crucial. “The Common Rule does not require consent for this activity” and “using the data this way is consistent with what we promised participants” are not interchangeable statements.

De-identification does not erase explicit promises

Imagine that participants were told, “Your data will only be used for research on cardiovascular disease.” Years later, researchers remove identifiers and want to use the same information for an unrelated purpose.

Whether the resulting dataset falls outside a particular regulatory definition does not retroactively change what participants were promised. An explicit limitation may remain ethically important even when researchers can no longer identify individual participants.

This is especially relevant when assessing how far secondary research can move from the originally authorized purpose. De-identification should not be used as a convenient mechanism for circumventing an explicit restriction.

Removing identifiers does not necessarily make re-identification impossible

Datasets can contain combinations of attributes that make individuals distinguishable even after obvious identifiers are removed. Exact dates, rare diagnoses, detailed geography, occupation, demographic characteristics, unusual events, and longitudinal patterns can become identifying when combined.

The surrounding information environment also changes. A dataset considered difficult to re-identify when it was released may become more revealing when new public databases, commercial datasets, linkage techniques, or computational methods become available.

Researchers should therefore avoid presenting de-identification as an absolute guarantee unless the applicable technical and legal standard genuinely supports that claim. More often, privacy protection is a matter of reducing identifiability and controlling access rather than making re-identification metaphysically impossible. The statisticians will forgive us for refusing to assign a probability of exactly zero.

Secondary use can create new information without identifying individuals

Ethical concerns are not limited to whether a researcher learns someone's name. Secondary analysis may generate sensitive inferences about groups, communities, or categories of people.

For example, de-identified data might be used to characterize a small ethnic community, geographic population, occupational group, school, or patient population in ways that create stigma or discrimination. No individual participant needs to be named for the research output to have consequences for people associated with that group.

These concerns vary considerably by research context, but they show why individual identifiability should not be treated as the sole measure of ethical risk.

The purpose of the secondary use still matters

Removing identifiers does not answer whether the proposed research purpose is scientifically defensible, proportionate, consistent with repository conditions, or otherwise appropriate.

The Philippine Data Privacy Act's implementing rules, for example, emphasize declared and legitimate purposes, compatible processing, proportionality, and appropriate safeguards when personal data are involved. They also contain specific provisions for data obtained from parties other than the data subject for research purposes.

Once information is genuinely outside the relevant definition of personal data, some privacy-law requirements may no longer apply in the same way. But researchers should establish that status under the applicable law rather than assuming that a home-grown de-identification procedure accomplished it.

Repository and contractual restrictions can survive de-identification

Data-use agreements may prohibit particular analyses, onward sharing, linkage, re-identification attempts, commercial uses, or transfer outside approved environments. These restrictions do not necessarily disappear when a researcher strips direct identifiers from a working copy.

The same applies to ethical responsibilities attached to archived research data. Access conditions and institutional commitments can govern the use of a dataset independently of whether individual records are readily identifiable.

De-identification should be designed before data reach the secondary researcher when possible

There is an important difference between receiving data from which you cannot readily ascertain identities and receiving fully identifiable records and personally removing the identifiers.

OHRP specifically notes that research can still involve human subjects when an investigator obtains information through interaction or intervention and then removes identifiers, because the investigator may already know or readily ascertain the participants' identities. For secondary research, what investigators obtain and whether they can readily ascertain identities are central considerations.

Where scientifically feasible, an authorized data custodian or intermediary can sometimes prepare the dataset so that the secondary investigator never receives identifying information in the first place.

Watch Out

Do not promise that a dataset is “anonymous” simply because names and identification numbers have been removed. Establish what identifiability standard applies, what information remains, who possesses linkage keys, and whether other available data could permit identities to be readily ascertained.

04 · A Practical Example

When Removing Names Changes the Regulation but Not the Promise

Hypothetical Example

A dataset was collected for a narrowly stated research purpose

Participants in an earlier study were explicitly told that their identifiable information would be used only for research concerning a particular health condition. A new team wants to study an unrelated sensitive topic. The data custodian proposes removing identifiers before releasing the dataset.

Assess identifiability The institution determines whether the resulting information would genuinely be non-identifiable to the secondary investigators under the applicable framework rather than assuming that deleting names is sufficient.
Assess regulatory status If investigators cannot readily ascertain participants' identities, the project may be treated differently under applicable human-subjects regulations.
Assess the original commitment The institution separately asks whether the proposed use conflicts with the explicit limitation participants were given.
Assess remaining risks The researchers consider whether the analysis could expose small groups, create sensitive inferences, enable linkage, or otherwise produce harms despite individual de-identification.
Decision The research does not become ethically acceptable merely because the regulatory identifiability question changed. The original promise and other applicable restrictions still need to be addressed.
05 · What Researchers Often Get Wrong

Common Misunderstandings About De-Identification and Ethics

Misconception

“No names means anonymous data.”

Removing direct identifiers is only one step. Other attributes, combinations of variables, code keys, or external information may still allow identities to be readily ascertained.

Misconception

“If the Common Rule does not apply, there is no ethical issue.”

Regulatory classification and ethical permissibility are not identical. Consent commitments, repository conditions, privacy obligations, contractual restrictions, scientific validity, and potential group harms may remain relevant.

Misconception

“De-identification cancels restrictions in the original consent.”

Not necessarily. OHRP advisory material specifically recognizes that institutions may still have obligations to honor agreements with participants even when secondary use of non-readily-identifiable material falls outside Common Rule human-subjects requirements.

Misconception

“De-identified data cannot harm anyone.”

Individual privacy risk may be reduced, but research can still affect communities or identifiable groups, generate stigmatizing conclusions, or violate expectations and institutional commitments.

Misconception

“I can receive identifiable data first and simply anonymize them myself.”

Receiving and handling identifiable private information may itself trigger obligations. If identifiers are unnecessary, consider whether an authorized custodian can prepare the data before the secondary research team receives them.

06 · What This Means for You

Use De-Identification as a Safeguard, Not as an Ethical Shortcut

When planning secondary research, reducing identifiability is often good practice. It can lower privacy risk and may substantially change regulatory requirements. But the analysis should not stop there.

A simple decision framework

If identifiers are unnecessary for the secondary study
Minimize them and consider having an authorized custodian prepare the data before release to the research team.
If a code key or combinations of variables could permit identification
Determine how the applicable framework classifies the information and restrict access accordingly.
If the original consent or data-use agreement restricts the proposed purpose
Do not assume de-identification overrides that restriction. Seek the appropriate ethics, legal, repository, or institutional determination.
If the analysis could create sensitive group-level conclusions
Assess those risks even when no individual participant can readily be identified.
07 · A Quick Checklist

Before Treating De-Identification as Sufficient for Secondary Use

Check more than whether names were removed:
Identify the formal definition of identifiable, coded, de-identified, or anonymous information that applies to your study.
Determine whether any researcher can access a code key or otherwise readily ascertain participants' identities.
Review combinations of demographic, geographic, temporal, clinical, or other variables that may increase re-identification risk.
Review the original consent for explicit restrictions or promises concerning future use.
Check repository licenses, data-use agreements, contractual restrictions, and institutional policies.
Assess whether the secondary analysis could create meaningful harms to identifiable groups or communities.
Prohibit unauthorized attempts to re-identify participants or link records back to identities.
Obtain the required institutional determination rather than independently declaring the research outside human-subjects requirements.
08 · Frequently Asked Questions

Frequently Asked Questions About De-Identification and Secondary Research

Is removing names enough to de-identify research data?

Not necessarily. Other information may permit identities to be readily ascertained, and a code key may still connect records to individuals. Apply the definition and technical standard relevant to your jurisdiction and research setting.

Are coded data the same as anonymous data?

No. Coded data have identifiers replaced with a code while a key exists that can reconnect the code to the individual. Whether coded data are considered identifiable to a particular researcher depends on access arrangements and the applicable regulatory framework.

Does research using de-identified data need ethics review?

The answer varies by jurisdiction and institution. Under the U.S. Common Rule, some secondary research involving information that investigators cannot readily link to individuals may not involve human subjects, but OHRP recommends that institutions designate an authorized person or entity to make such determinations.

Can de-identified data be used for a purpose participants explicitly prohibited?

Do not assume so. Even when a secondary use falls outside a particular human-subjects regulation, commitments made to participants may remain ethically or institutionally binding.

Can de-identified data become identifiable again?

Potentially. Detailed records, linkage keys, external datasets, and new analytical techniques can increase re-identification possibilities. The actual risk depends on the information retained and the surrounding data environment.

Does de-identification mean researchers can publish any subgroup analysis?

No. Very small or distinctive groups may become recognizable, and findings may create stigma or other harms even without identifying individual participants. Disclosure and group-level risks should be considered separately.

09 · The Bottom Line

De-Identification Changes the Privacy Question, Not Every Ethical Question

The Bottom Line

De-identification can reduce privacy risks and may change the regulatory status of secondary research, but it does not automatically make the proposed use ethically acceptable.

Verify what “de-identified” actually means in your setting, then separately consider original commitments, purpose, access restrictions, residual re-identification risk, group harms, and applicable governance. Removing identity from the dataset does not remove the dataset's ethical history.

10 · Sources and Further Reading

Authoritative Guidance on De-Identification and Secondary Research

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes