03 · What You Need to Know
Identifiability Is One Part of the Ethical Analysis, Not the Whole of It
De-identification can materially change the regulatory analysis
Identifiability matters because many human-subjects and privacy frameworks impose different requirements depending on whether researchers can identify the people behind the information.
Under the U.S. Common Rule, for example, a human subject includes a living individual about whom an investigator obtains, uses, studies, analyzes, or generates identifiable private information or identifiable biospecimens. OHRP states that secondary research involving only coded private information or coded biospecimens may not involve human subjects when investigators cannot readily ascertain the identities of the individuals because access to the code key is appropriately restricted.
The Common Rule also contains an exemption for certain secondary research in which information is recorded so that subjects' identities cannot readily be ascertained directly or through linked identifiers, provided that investigators do not contact or re-identify subjects.
De-identified, coded, and anonymized do not necessarily mean the same thing
Researchers often use these terms loosely, but regulatory definitions can differ. A coded dataset may have names replaced by numbers while a separate key still connects those numbers to individuals. Whether that dataset is considered identifiable can depend on who possesses the key, whether the research team can obtain it, and which regulatory framework applies.
Direct identifiers removed
Names or obvious identifiers have been removed, but other variables or a code key may still permit identification.
Coded data
Identifiers have been replaced by a code and a key exists somewhere that can reconnect the code to the individual.
Non-identifiable to the investigator
The investigator cannot readily ascertain identities directly, through a coding system, or through other reasonably available means under the applicable standard.
Do not decide which category applies merely by looking at whether the spreadsheet contains a “Name” column. The actual data, access arrangements, available external information, and governing definitions matter.
Regulatory status and ethical acceptability are different questions
A study can fall outside a particular human-subjects regulation without becoming ethically neutral. OHRP's own advisory material illustrates the distinction clearly in the context of biospecimens: a secondary use of non-readily-identifiable material may fall outside Common Rule human-subjects requirements, while researchers and institutions may still have an obligation to honor agreements they made with participants about how their material would be used.
This distinction is crucial. “The Common Rule does not require consent for this activity” and “using the data this way is consistent with what we promised participants” are not interchangeable statements.
De-identification does not erase explicit promises
Imagine that participants were told, “Your data will only be used for research on cardiovascular disease.” Years later, researchers remove identifiers and want to use the same information for an unrelated purpose.
Whether the resulting dataset falls outside a particular regulatory definition does not retroactively change what participants were promised. An explicit limitation may remain ethically important even when researchers can no longer identify individual participants.
This is especially relevant when assessing how far secondary research can move from the originally authorized purpose. De-identification should not be used as a convenient mechanism for circumventing an explicit restriction.
Removing identifiers does not necessarily make re-identification impossible
Datasets can contain combinations of attributes that make individuals distinguishable even after obvious identifiers are removed. Exact dates, rare diagnoses, detailed geography, occupation, demographic characteristics, unusual events, and longitudinal patterns can become identifying when combined.
The surrounding information environment also changes. A dataset considered difficult to re-identify when it was released may become more revealing when new public databases, commercial datasets, linkage techniques, or computational methods become available.
Researchers should therefore avoid presenting de-identification as an absolute guarantee unless the applicable technical and legal standard genuinely supports that claim. More often, privacy protection is a matter of reducing identifiability and controlling access rather than making re-identification metaphysically impossible. The statisticians will forgive us for refusing to assign a probability of exactly zero.
Secondary use can create new information without identifying individuals
Ethical concerns are not limited to whether a researcher learns someone's name. Secondary analysis may generate sensitive inferences about groups, communities, or categories of people.
For example, de-identified data might be used to characterize a small ethnic community, geographic population, occupational group, school, or patient population in ways that create stigma or discrimination. No individual participant needs to be named for the research output to have consequences for people associated with that group.
These concerns vary considerably by research context, but they show why individual identifiability should not be treated as the sole measure of ethical risk.
The purpose of the secondary use still matters
Removing identifiers does not answer whether the proposed research purpose is scientifically defensible, proportionate, consistent with repository conditions, or otherwise appropriate.
The Philippine Data Privacy Act's implementing rules, for example, emphasize declared and legitimate purposes, compatible processing, proportionality, and appropriate safeguards when personal data are involved. They also contain specific provisions for data obtained from parties other than the data subject for research purposes.
Once information is genuinely outside the relevant definition of personal data, some privacy-law requirements may no longer apply in the same way. But researchers should establish that status under the applicable law rather than assuming that a home-grown de-identification procedure accomplished it.
Repository and contractual restrictions can survive de-identification
Data-use agreements may prohibit particular analyses, onward sharing, linkage, re-identification attempts, commercial uses, or transfer outside approved environments. These restrictions do not necessarily disappear when a researcher strips direct identifiers from a working copy.
The same applies to ethical responsibilities attached to archived research data. Access conditions and institutional commitments can govern the use of a dataset independently of whether individual records are readily identifiable.
De-identification should be designed before data reach the secondary researcher when possible
There is an important difference between receiving data from which you cannot readily ascertain identities and receiving fully identifiable records and personally removing the identifiers.
OHRP specifically notes that research can still involve human subjects when an investigator obtains information through interaction or intervention and then removes identifiers, because the investigator may already know or readily ascertain the participants' identities. For secondary research, what investigators obtain and whether they can readily ascertain identities are central considerations.
Where scientifically feasible, an authorized data custodian or intermediary can sometimes prepare the dataset so that the secondary investigator never receives identifying information in the first place.
Watch Out
Do not promise that a dataset is “anonymous” simply because names and identification numbers have been removed. Establish what identifiability standard applies, what information remains, who possesses linkage keys, and whether other available data could permit identities to be readily ascertained.
07 · A Quick Checklist
Before Treating De-Identification as Sufficient for Secondary Use
Check more than whether names were removed:
Identify the formal definition of identifiable, coded, de-identified, or anonymous information that applies to your study.
Determine whether any researcher can access a code key or otherwise readily ascertain participants' identities.
Review combinations of demographic, geographic, temporal, clinical, or other variables that may increase re-identification risk.
Review the original consent for explicit restrictions or promises concerning future use.
Check repository licenses, data-use agreements, contractual restrictions, and institutional policies.
Assess whether the secondary analysis could create meaningful harms to identifiable groups or communities.
Prohibit unauthorized attempts to re-identify participants or link records back to identities.
Obtain the required institutional determination rather than independently declaring the research outside human-subjects requirements.