01 · The Question
When Should Researchers Break the Link to Participant Identity for Good?
A research team finishes data collection. Names and contact information have already been separated from the analysis dataset, but a protected key can still reconnect participant codes to identities.
Should the team destroy that key immediately?
Perhaps, but permanent removal is more consequential than routine separation. Once the only viable identity link is destroyed, researchers may lose the ability to contact participants, honour some withdrawal requests, perform participant-level follow-up, verify particular records, link future data, or carry out other approved functions.
The right time is therefore not simply "when data collection ends." It is when continued identifiability no longer serves a justified purpose and permanent removal is consistent with the requirements governing the research.
03 · What You Need to Know
Permanent Removal Is a Research Decision, Not Just a Data-Cleaning Step
Separation and Permanent Removal Solve Different Problems
Researchers should first distinguish two actions that are easily conflated.
Separating identifiers
Direct identifiers are kept apart from substantive research data, while authorised re-identification remains possible through a code key or other linkage mechanism.
Permanently removing the identity link
The relevant identifiers, linkage information, or other re-identification mechanism are eliminated so that the research team can no longer restore identity through that route.
Researchers can therefore separate participant identifiers from research data long before they are ready to destroy the ability to reconnect them.
That distinction allows a study to reduce routine identity exposure while preserving participant-specific functions that are still necessary.
Do Not Destroy the Link While the Study Still Needs Participant-Specific Follow-Up
Longitudinal studies provide the clearest example. If researchers need to invite participants to later waves, connect repeated measurements, schedule assessments, or perform other participant-specific procedures, some controlled identity mechanism may still be necessary.
Likewise, intervention or clinical research may require participant-specific safety procedures or follow-up. Other studies may require approved record linkage or additional data collection.
Permanent removal before those functions are complete could make the study impossible to conduct as approved.
Withdrawal Can Depend on Whether Participants Can Still Be Located in the Dataset
Researchers should also consider what they have told participants about withdrawal.
If a participant asks researchers to remove their data, the team may need an identity link to locate that person's records. Once data have been effectively anonymized and the connection has been irreversibly broken, locating an individual's contribution may no longer be possible.
This does not mean researchers must preserve identifiers indefinitely merely to make withdrawal technically possible. The relevant ethics, legal, institutional, and consent requirements should determine what withdrawal rights or options apply and how they should be described.
The practical point is that irreversible anonymization changes what researchers can do later. Participants should not be promised individual data removal after anonymization if the research architecture makes that impossible.
Verification and Research Integrity May Justify Retention
Research data can have purposes beyond the immediate statistical analysis. Records may need to support verification, quality assurance, audit, reproducibility, investigation of research integrity concerns, regulatory obligations, or other legitimate research functions.
Whether those purposes require identifiable information rather than coded or less identifying records should be assessed rather than assumed.
UKRI's research-data retention framework explicitly supports risk-proportionate decisions about retaining or potentially destroying research data and records and emphasizes documenting those decisions. It is not a universal retention schedule for all research, but it illustrates why destruction should be governed rather than improvised.
Data Protection Does Not Necessarily Require Researchers to Delete Personal Data as Soon as a Project Ends
A common misconception is that data-protection law imposes a universal requirement to delete identifiable research information immediately after publication or project completion.
Under the UK GDPR research provisions, for example, personal data can be retained for longer periods, potentially indefinitely, when processed solely for qualifying research-related purposes and appropriate safeguards are in place. The ICO specifically notes that the research exception modifies the ordinary storage-limitation principle.
This is jurisdiction-specific. Researchers should not transplant UK rules into another country. The broader lesson is that "the study is finished" and "identifiable information must now be deleted" are not universally equivalent propositions.
Retention Still Needs a Purpose
Permission to retain research information is not the same as a reason to retain every identifier forever.
The ICO's general storage-limitation guidance requires organisations to justify how long personal data are retained, periodically review them, and erase or anonymize information that is no longer needed, while recognising special provisions for research and archiving.
A useful review therefore separates two questions:
- Do we still need the research data?
- Do we still need the data to remain identifiable?
The answers may differ. Researchers may legitimately retain the scientific dataset while eliminating names, contact details, or the linkage mechanism when those no longer serve the continuing research purpose.
Destroying the Code Key May Be the Critical Step, but Not Always the Only Step
Suppose an analysis dataset contains study codes and a separate file connects those codes to names. Permanently destroying the linking file removes an obvious re-identification route.
That does not automatically mean the remaining dataset is anonymous.
Exact dates, rare occupations, detailed geography, unusual demographic combinations, free-text responses, images, genomic information, or other variables may still make participants identifiable. Researchers therefore need to assess whether the remaining data are actually anonymous rather than treating destruction of the key as sufficient proof.
Watch Out
"We destroyed the code key" and "the dataset is anonymous" are not necessarily the same statement. The remaining information must still be assessed for direct, indirect, contextual, and linkage-based identification.
Different Identifiers Can Be Removed at Different Times
Researchers do not need to treat all identifying information as one package.
A postal address needed to send study materials may become unnecessary shortly after enrolment. A telephone number may remain useful through active follow-up. A code key may remain necessary through longitudinal analysis. Signed consent documentation may be subject to separate retention requirements.
This creates an identifier lifecycle rather than one project-wide deletion date.
| Information |
Possible Continuing Purpose |
Question Before Permanent Removal |
| Contact information |
Follow-up, scheduling, participant communication |
Will any authorised participant contact still occur? |
| Code key |
Longitudinal linkage, withdrawal, verification |
Does any approved function still require re-identification? |
| Direct identifiers in working data |
Sometimes none after early study stages |
Can they be removed while the necessary linkage is preserved separately? |
| Consent documentation |
Documentation and regulatory or institutional requirements |
What retention requirements govern these records? |
| Identifiable source records |
Audit, verification, integrity, future approved research |
Does continued identifiability remain necessary for those purposes? |
Future Research Is Not Automatically a Reason to Keep Everything Identifiable
Researchers may reasonably anticipate future research uses, but "someone might want this later" is not a complete retention strategy.
Future use should be considered alongside the permissions, lawful basis, participant information, governance arrangements, institutional policy, funder or sponsor requirements, and scientific need applicable to the dataset.
Where future research does not require identities, anonymized or less identifiable forms may preserve much of the scientific value with lower confidentiality risk. Where legitimate future research requires person-level linkage, continued pseudonymization with safeguards may be more appropriate.
Anonymization Can Change the Regulatory Status of Information
Under data-protection regimes such as the UK GDPR, effectively anonymized information falls outside the personal-data regime because it no longer relates to an identified or identifiable person under the applicable standard. Pseudonymized data generally remain personal data.
UKRI's updated research guidance accordingly distinguishes anonymisation from pseudonymisation and emphasizes that identifiability depends on both content and context.
Researchers should therefore document what they mean by "removing identifiers." Deleting a name field, destroying a key, and achieving effective anonymization are different events.
Permanent Removal Should Be Documented
Irreversible destruction can affect future research possibilities and participant rights or expectations. The decision should therefore be traceable.
Depending on the study, documentation might record what information was removed, when it occurred, which linkage mechanisms were destroyed, what information was retained, why the action was taken, and under which retention or data-management plan.
This is particularly useful when datasets outlive individual researchers. A future custodian should not have to infer from an abandoned folder named "FINAL_FINAL_ANON_v3" what was actually done.
07 · A Quick Checklist
Before Permanently Removing Participant Identifiers
Before breaking the identity link permanently, check:
Confirm that participant-specific follow-up, communication, or study procedures no longer require identification.
Determine whether approved longitudinal or cross-dataset linkage still requires the code key or other identifiers.
Check what participants were told about withdrawal and whether permanent anonymization changes the ability to locate individual records.
Review legal, regulatory, institutional, sponsor, funder, ethics, research-integrity, and retention requirements before destruction.
Distinguish information that must be retained from information that must remain identifiable.
Determine whether different identifiers can be removed at different stages rather than applying one date to all records.
Assess whether the remaining dataset could still identify participants after the direct linkage mechanism is destroyed.
Document what was permanently removed, when, why, and under which approved retention or data-management arrangement.