01 · The Question
Can Two Existing Datasets Be Combined Without Asking Everyone Again?
A researcher has one dataset containing survey responses and another containing clinical, educational, administrative, or other records. Linking the two could answer a question that neither dataset can answer alone.
But linkage requires matching records that belong to the same person, which may involve names, identification numbers, dates of birth, addresses, or other identifying information at some stage. Does that mean every participant must be contacted again before the datasets can be connected?
03 · What You Need to Know
Permission to Use Two Datasets Separately Does Not Automatically Include Permission to Link Them
Data linkage creates a new research resource
Data linkage, sometimes called record linkage, connects information about the same individual across two or more data sources. Researchers might link a longitudinal survey to hospital records, educational records to employment information, registry data to mortality records, or research data to administrative databases.
The scientific advantage is obvious: information distributed across separate systems can be studied together. But the linked dataset may also reveal relationships and characteristics that neither source exposed on its own.
This is why the ethical question is not merely whether dataset A is legitimate and dataset B is legitimate. You also need to ask whether A plus B is legitimate.
Linkage may require identifiers even when the final research file does not
To determine that two records belong to the same person, someone usually needs a way to match them. Exact linkage may use a common unique identifier. Probabilistic linkage may use combinations of information such as names, dates of birth, addresses, sex, or other attributes to estimate whether records correspond to the same individual.
The final analytical dataset can nevertheless be designed so that researchers never receive those direct identifiers. Identifying information can be handled by an authorized linkage unit, data custodian, or other intermediary that performs the match and then supplies researchers with appropriately coded or non-identifiable research data.
Linkage identifiers
Information needed by an authorized party to determine which records belong to the same individual.
Analysis data
The variables researchers need to answer the research question after linkage has been completed.
Separating these functions can reduce the number of people who handle identifying information and may materially alter the privacy and regulatory analysis.
Recontact depends partly on what participants originally authorized
If participants were explicitly told that their research information could be linked with specified records or categories of records for future research, a proposed linkage may already fall within the authorization, subject to applicable review and governance requirements.
If linkage was never mentioned, the answer becomes more contextual. Researchers need to assess what the original consent permits when it does not clearly cover the new study and whether another lawful and ethically permissible pathway is available.
An explicit refusal of linkage should not be treated as equivalent to silence.
Consent is not always required for every secondary linkage study
Some research frameworks permit secondary research without obtaining new individual consent when specified conditions are satisfied. Under the U.S. Common Rule, for example, certain secondary research uses of identifiable private information may qualify for exemption under 45 CFR 46.104(d)(4), while nonexempt research may sometimes proceed under an IRB-approved waiver of consent when the criteria in 45 CFR 46.116 are met.
OHRP also distinguishes secondary research in which investigators obtain identifiable private information from research in which investigators cannot readily ascertain individuals' identities. The latter may fall outside the Common Rule definition of human-subjects research under specified circumstances.
These pathways are context-specific. They should not be converted into a general rule that “data linkage does not require consent.”
Recontact itself can change the research
Requiring individual consent for linkage can have scientific consequences. Some individuals may be unreachable, some may decline, and willingness to consent may differ systematically across groups. The resulting linked sample may therefore differ from the population researchers intended to study.
Empirical research on record linkage has documented wide variation in consent proportions and has raised concerns about selection bias when linkage is limited to those who provide specific consent. That does not make consent ethically dispensable. It does mean that the consequences of requiring recontact can be relevant when an ethics committee assesses whether a waiver is justified under an applicable framework.
“Impracticable” does not simply mean expensive or inconvenient
Large administrative datasets may contain millions of records, and current contact details may not exist. Those circumstances can be relevant to waiver decisions. Researchers should nevertheless avoid treating scale as an automatic exemption from consent.
Where a waiver is sought, the reviewing body needs the information required by the applicable standard, which may include the study's risk, practicability without the waiver, effects on participants' rights and welfare, and the necessity of using identifiable information.
The broader principles governing when secondary research can proceed without new consent apply to linkage as well.
Who performs the linkage can be as important as whether linkage occurs
A strong privacy-preserving design may separate identifiers from research variables. For example, data custodians can send identifying fields to an authorized linkage unit, while researchers receive only the linked analytical variables with a project-specific identifier.
OHRP materials use the concept of an “honest broker” for a neutral intermediary between individuals whose data or tissue are studied and the researcher. Similar separation principles can reduce researchers' direct access to identifying information.
This question becomes especially important when deciding who should be allowed to perform linkage using identifiable data.
Linkage can make previously low-risk information more revealing
Imagine one dataset contains a person's employment history and another contains health information. Separately, each dataset exposes only part of the person's circumstances. Linking them may reveal associations between workplace, diagnosis, treatment, absence, income, or other characteristics.
In other words, privacy risk is not simply the sum of the risks in the source datasets. Linkage can generate new information and make individuals more distinguishable. The question of whether linkage creates new privacy risks from otherwise safe datasets therefore requires separate attention.
Linkage quality is also an ethical issue
Record linkage is not always perfectly accurate. False matches connect records belonging to different people. Missed matches fail to connect records that belong to the same person. These errors can distort findings and may affect some demographic groups more than others when identifiers are incomplete, inconsistent, or differently formatted.
A responsible linkage protocol should therefore address linkage quality, validation, error rates where measurable, and the potential consequences of linkage error. Privacy protection is essential, but a beautifully secure linkage that systematically matches the wrong people is not a methodological triumph.
The linked file should contain only what the research needs
Linkage can encourage data accumulation because researchers suddenly have access to many variables across multiple systems. Resist the urge to retain everything simply because the infrastructure can combine it.
The Philippine Data Privacy Act's implementing rules articulate proportionality by requiring personal-data processing to be adequate, relevant, suitable, necessary, and not excessive in relation to the declared purpose. They also impose requirements on data sharing and further processing.
Where applicable, design the linkage around the minimum identifiers necessary for matching and the minimum analytical variables necessary for the approved research question.
Watch Out
Do not perform exploratory matching with identifiable records simply to see whether linkage “works” before obtaining the required authorization. The act of accessing, transferring, and matching identifying information can itself constitute regulated processing or human-subjects research activity.
07 · A Quick Checklist
Before Linking Existing Datasets Without Recontacting Participants
Before performing the linkage, check:
Define why linkage is necessary to answer the research question rather than merely useful because more data are available.
Review the consent, legal authority, repository conditions, and data-use agreements governing each source dataset.
Determine whether linkage itself falls within the authorized purposes of each dataset.
Obtain any required ethics, privacy, institutional, or data-custodian approval for linkage without recontact.
Identify the minimum fields needed to match records and the minimum variables needed for analysis.
Specify who may access direct identifiers and whether linkage can be performed by a separate authorized intermediary.
Document how identifiers, linkage keys, and analytical data will be separated, stored, transferred, and eventually destroyed or retained.
Evaluate false-match and missed-match risks and how linkage errors could affect the study's conclusions.
Reassess privacy and disclosure risk after linkage rather than assuming the linked dataset has the same risk profile as its sources.