03 · What You Need to Know
Linkage Changes What the Data Can Reveal
Privacy risk is not necessarily additive
If one dataset has a low privacy risk and another has a low privacy risk, it does not follow that the linked dataset has only a slightly larger risk. The relationship can be nonlinear.
Dataset A might contain age, occupation, and neighborhood. Dataset B might contain dates of service and a particular health condition. Neither contains a name. Once combined, however, the resulting profile may describe so distinctive a person that identification becomes considerably easier.
Linkage therefore changes the informational environment rather than merely placing additional columns beside existing ones.
Quasi-identifiers become more powerful in combination
Some variables are not direct identifiers but can help distinguish individuals. These are often called quasi-identifiers. Examples may include age, sex, postal area, occupation, dates, educational institution, household structure, or unusual events.
The problem is combinatorial. A single characteristic may describe thousands of people. Several characteristics taken together may describe very few.
Direct identifier
Information such as a name or another field that directly identifies a particular person.
Quasi-identifier
Information that may not identify someone alone but can contribute to identification when combined with other attributes or external information.
Consequently, a dataset that was reasonably protected before linkage may need stronger disclosure controls afterward.
Linkage can turn coded information back into identifiable information
A coded dataset contains a code in place of direct identifiers while a key exists that can reconnect the code with identifying information. Under OHRP guidance, whether coded private information is identifiable to a secondary investigator depends partly on whether that investigator can readily ascertain the person's identity, including through access to the coding system.
OHRP gives examples in which secondary investigators cannot readily ascertain identities because agreements, repository procedures, or legal requirements prohibit release of the code key. If those arrangements change and investigators can readily ascertain identities, the regulatory status of the research may change as well.
Linkage architecture therefore matters. A research team that never possesses the identity key occupies a different position from a team that receives identifiers, performs the match, and then deletes them afterward.
Linkage can reveal information that did not exist in either source alone
Suppose one dataset records employment and another records clinical encounters. Linking them may reveal patterns between workplace and diagnosis. Link educational records with household data and researchers may infer socioeconomic or family circumstances. Link location histories with service records and patterns of behavior may become visible.
These are not merely additional variables. Some information emerges from relationships among variables across sources.
This means that privacy assessment should consider derived and inferred information, not just the fields originally collected.
A safe release decision is contextual
A dataset is not simply “safe” or “unsafe” in the abstract. Risk depends on who receives it, what other information they can access, which linkage tools are available, what contractual restrictions exist, and what technical environment surrounds the data.
A coded research file supplied to investigators who are contractually and technically prevented from obtaining the key may create a different risk from the same file supplied to researchers who also possess a registry containing identities.
OHRP's guidance reflects this contextual approach by focusing on whether investigators can readily ascertain identity rather than treating coding as an intrinsic property that produces the same regulatory result in every situation.
Linkage can increase the sensitivity of otherwise ordinary information
Individual variables do not always appear sensitive in isolation. School attended, occupation, neighborhood, service utilization, income category, or attendance records may seem relatively ordinary in one context. Combined with another dataset, they may reveal health conditions, disability, family circumstances, migration status, financial hardship, or other sensitive information.
The privacy question should therefore include: what will someone know after linkage that they could not know before?
Linkage may create risks for people who never appear by name
Researchers often focus on direct re-identification, but linked data can also reveal characteristics of small groups, households, institutions, or communities.
A combination of geographic, demographic, clinical, or educational information may allow readers to recognize a small population even if individual records remain unnamed. Published tables, maps, case descriptions, or subgroup analyses can sometimes create disclosure risks that were not apparent in the analytical dataset itself.
Output checking should therefore be part of privacy protection rather than an afterthought performed just before manuscript submission.
More linkage can make later linkage easier
Once researchers create a richly linked dataset, it may become an attractive foundation for additional linkages. Each new source can increase informational depth and potentially create new routes to identification.
This is one reason authorization for one linkage should not casually be treated as authorization for an indefinite chain of future combinations. The question of whether datasets may be linked without recontacting participants should be assessed for the proposed linkage rather than assumed from the mere existence of an earlier linked resource.
Security controls should reflect the linked dataset, not its least sensitive source
The Philippine Data Privacy Act's implementing rules require reasonable and appropriate organizational, physical, and technical security measures for personal data and require protection against unlawful access, disclosure, alteration, destruction, and other unlawful processing.
The same rules emphasize proportionality, requiring personal-data processing to be adequate, relevant, necessary, and not excessive for the declared purpose.
For linked research, that means security and access controls should be designed around the resulting resource and its risks. If linkage produces a substantially richer or more sensitive dataset, protections suitable for one original source may no longer be adequate.
Linkage errors create another kind of privacy problem
False matches do more than distort statistical estimates. They can attribute one person's information to another person. If linked records are subsequently used for participant contact, individual reporting, decision-making, or other person-level activities, such errors can have particularly serious consequences.
Researchers should therefore consider linkage accuracy alongside confidentiality. The strongest privacy controls in the server room cannot rescue a study that has confidently assembled the wrong person's record.
Privacy-preserving architecture can reduce unnecessary exposure
One approach is functional separation. The party performing linkage receives the minimum identifying information needed for matching, while analytical researchers receive only approved linked variables with project-specific codes.
OHRP materials describe an “honest broker” as a neutral intermediary between the individual whose data or tissue are studied and the researcher. Comparable intermediary models can help keep identifying information away from researchers who do not need it.
The related question of who should be permitted to perform linkage using identifiable information is therefore part of privacy design, not merely project administration.
Watch Out
Do not assess the privacy of a linked dataset by checking only whether either source contains names. Ask what combinations, inferences, code keys, external sources, and future linkages could allow someone with access to learn or identify after the datasets are combined.