Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Does Data Linkage Create New Privacy Risks Even When Each Dataset Was Safe on Its Own?

Two datasets that present modest privacy risks separately can become substantially more revealing when linked. Researchers should reassess identifiability, sensitivity, access, disclosure, and security after linkage rather than relying on the protections applied to each source alone.

372
Privacy Risks Created by Data Linkage Guide 372 of 398
01 · The Question

Can Combining Two Safe Datasets Create an Unsafe One?

Dataset A contains information that seems reasonably protected. Dataset B does too. Neither appears especially identifying on its own. It is tempting to conclude that linking them simply produces a larger dataset with roughly the same privacy profile.

That conclusion can be wrong. Linkage may create unique combinations of attributes, reveal relationships that neither source disclosed independently, reconnect coded information with identities, or make individuals easier to distinguish. Privacy risk therefore needs to be assessed again after the datasets are combined.

02 · The Short Answer

Privacy Risk Can Increase When Datasets Are Combined

In Brief

Yes. Linking datasets can create new privacy risks even when each source was reasonably safe on its own because the combined information may make people more identifiable, reveal new sensitive facts, or enable inferences that neither dataset supported separately.

The linked dataset should therefore receive its own privacy and security assessment. Researchers need to consider the information created by the combination, who can access linkage identifiers and keys, whether the resulting records remain appropriately protected, and how outputs could disclose individuals or groups.

03 · What You Need to Know

Linkage Changes What the Data Can Reveal

Privacy risk is not necessarily additive

If one dataset has a low privacy risk and another has a low privacy risk, it does not follow that the linked dataset has only a slightly larger risk. The relationship can be nonlinear.

Dataset A might contain age, occupation, and neighborhood. Dataset B might contain dates of service and a particular health condition. Neither contains a name. Once combined, however, the resulting profile may describe so distinctive a person that identification becomes considerably easier.

Linkage therefore changes the informational environment rather than merely placing additional columns beside existing ones.

Quasi-identifiers become more powerful in combination

Some variables are not direct identifiers but can help distinguish individuals. These are often called quasi-identifiers. Examples may include age, sex, postal area, occupation, dates, educational institution, household structure, or unusual events.

The problem is combinatorial. A single characteristic may describe thousands of people. Several characteristics taken together may describe very few.

Direct identifier Information such as a name or another field that directly identifies a particular person.
Quasi-identifier Information that may not identify someone alone but can contribute to identification when combined with other attributes or external information.

Consequently, a dataset that was reasonably protected before linkage may need stronger disclosure controls afterward.

Linkage can turn coded information back into identifiable information

A coded dataset contains a code in place of direct identifiers while a key exists that can reconnect the code with identifying information. Under OHRP guidance, whether coded private information is identifiable to a secondary investigator depends partly on whether that investigator can readily ascertain the person's identity, including through access to the coding system.

OHRP gives examples in which secondary investigators cannot readily ascertain identities because agreements, repository procedures, or legal requirements prohibit release of the code key. If those arrangements change and investigators can readily ascertain identities, the regulatory status of the research may change as well.

Linkage architecture therefore matters. A research team that never possesses the identity key occupies a different position from a team that receives identifiers, performs the match, and then deletes them afterward.

Linkage can reveal information that did not exist in either source alone

Suppose one dataset records employment and another records clinical encounters. Linking them may reveal patterns between workplace and diagnosis. Link educational records with household data and researchers may infer socioeconomic or family circumstances. Link location histories with service records and patterns of behavior may become visible.

These are not merely additional variables. Some information emerges from relationships among variables across sources.

This means that privacy assessment should consider derived and inferred information, not just the fields originally collected.

A safe release decision is contextual

A dataset is not simply “safe” or “unsafe” in the abstract. Risk depends on who receives it, what other information they can access, which linkage tools are available, what contractual restrictions exist, and what technical environment surrounds the data.

A coded research file supplied to investigators who are contractually and technically prevented from obtaining the key may create a different risk from the same file supplied to researchers who also possess a registry containing identities.

OHRP's guidance reflects this contextual approach by focusing on whether investigators can readily ascertain identity rather than treating coding as an intrinsic property that produces the same regulatory result in every situation.

Linkage can increase the sensitivity of otherwise ordinary information

Individual variables do not always appear sensitive in isolation. School attended, occupation, neighborhood, service utilization, income category, or attendance records may seem relatively ordinary in one context. Combined with another dataset, they may reveal health conditions, disability, family circumstances, migration status, financial hardship, or other sensitive information.

The privacy question should therefore include: what will someone know after linkage that they could not know before?

Linkage may create risks for people who never appear by name

Researchers often focus on direct re-identification, but linked data can also reveal characteristics of small groups, households, institutions, or communities.

A combination of geographic, demographic, clinical, or educational information may allow readers to recognize a small population even if individual records remain unnamed. Published tables, maps, case descriptions, or subgroup analyses can sometimes create disclosure risks that were not apparent in the analytical dataset itself.

Output checking should therefore be part of privacy protection rather than an afterthought performed just before manuscript submission.

More linkage can make later linkage easier

Once researchers create a richly linked dataset, it may become an attractive foundation for additional linkages. Each new source can increase informational depth and potentially create new routes to identification.

This is one reason authorization for one linkage should not casually be treated as authorization for an indefinite chain of future combinations. The question of whether datasets may be linked without recontacting participants should be assessed for the proposed linkage rather than assumed from the mere existence of an earlier linked resource.

Security controls should reflect the linked dataset, not its least sensitive source

The Philippine Data Privacy Act's implementing rules require reasonable and appropriate organizational, physical, and technical security measures for personal data and require protection against unlawful access, disclosure, alteration, destruction, and other unlawful processing.

The same rules emphasize proportionality, requiring personal-data processing to be adequate, relevant, necessary, and not excessive for the declared purpose.

For linked research, that means security and access controls should be designed around the resulting resource and its risks. If linkage produces a substantially richer or more sensitive dataset, protections suitable for one original source may no longer be adequate.

Linkage errors create another kind of privacy problem

False matches do more than distort statistical estimates. They can attribute one person's information to another person. If linked records are subsequently used for participant contact, individual reporting, decision-making, or other person-level activities, such errors can have particularly serious consequences.

Researchers should therefore consider linkage accuracy alongside confidentiality. The strongest privacy controls in the server room cannot rescue a study that has confidently assembled the wrong person's record.

Privacy-preserving architecture can reduce unnecessary exposure

One approach is functional separation. The party performing linkage receives the minimum identifying information needed for matching, while analytical researchers receive only approved linked variables with project-specific codes.

OHRP materials describe an “honest broker” as a neutral intermediary between the individual whose data or tissue are studied and the researcher. Comparable intermediary models can help keep identifying information away from researchers who do not need it.

The related question of who should be permitted to perform linkage using identifiable information is therefore part of privacy design, not merely project administration.

Watch Out

Do not assess the privacy of a linked dataset by checking only whether either source contains names. Ask what combinations, inferences, code keys, external sources, and future linkages could allow someone with access to learn or identify after the datasets are combined.

04 · A Practical Example

Two Low-Risk Files Become Much More Revealing Together

Hypothetical Example

Linking an employee survey with health-service records

A research team receives two separately controlled datasets. The first contains employee survey information, including age band, work unit, occupation, and general well-being measures. The second contains health-service utilization data with dates and broad diagnostic categories. Neither dataset contains employee names in the research copy.

Before linkage Each dataset contains limited information. Individual employees may be difficult for the analytical researchers to distinguish confidently.
After linkage Occupation, small work unit, age band, unusual service dates, and diagnostic information now appear in the same record. Some combinations are much more distinctive.
New inference The linked information may reveal health characteristics associated with particular employees or small occupational groups that neither source disclosed alone.
Response The research team reassesses disclosure risk, limits variables, separates the linkage key, restricts access, applies output controls, and reviews whether subgroup reporting could reveal individuals or very small groups.

The datasets did not become ethically problematic merely because they were linked. Rather, linkage changed what could be learned from them, which changed the risk assessment and the safeguards needed.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Privacy After Data Linkage

Misconception

“Both datasets were de-identified, so the linked dataset is de-identified too.”

Not automatically. Linking may create distinctive combinations of variables or reconnect information through codes and external sources. Identifiability should be reassessed after linkage.

Misconception

“Linkage only adds variables; it does not create information.”

Relationships across datasets can support new inferences that neither source allowed independently. The informational content of the linked resource can therefore exceed the simple sum of its columns.

Misconception

“If researchers promise not to look up identities, identifiability is irrelevant.”

Intent matters less than actual access and capability under many regulatory definitions. If investigators can readily ascertain identities through a key or other means, promising not to do so may not make the information non-identifiable.

Misconception

“Privacy risk ends when the statistical analysis is finished.”

Research outputs can disclose sensitive information through small cells, detailed maps, unusual combinations, quotations, or subgroup descriptions. Disclosure review should extend to dissemination.

Misconception

“The security rules for the original datasets are automatically sufficient.”

A linked dataset may be more sensitive or identifying than either source. Security, access, retention, and disclosure controls should reflect the risk of the resulting linked resource.

06 · What This Means for You

Reassess Privacy After Linkage, Not Just Before It

Before linking data, assess each source. After linkage, assess the resulting dataset again. The second assessment is not redundant because the informational content, identifiability, and consequences may have changed.

A simple decision framework

If linkage creates new combinations of quasi-identifiers
Reassess re-identification and disclosure risk rather than inheriting the classification of the original files.
If researchers do not need identities
Separate linkage from analysis and prevent analytical researchers from accessing identifiers or the linkage key where feasible.
If linkage produces substantially more sensitive information
Strengthen access, security, retention, and output controls to match the resulting risk.
If additional linkage is later proposed
Assess the new combination again rather than assuming approval or privacy analysis for the first linkage automatically extends to the next.
07 · A Quick Checklist

Before Treating a Linked Dataset as Privacy-Safe

After linkage, check:
Identify what new information or inferences become possible only because the datasets were combined.
Reassess whether combinations of variables make individuals more readily distinguishable or identifiable.
Determine who possesses identifiers, matching information, and linkage keys and whether each person genuinely needs that access.
Remove variables that are unnecessary for the approved research question where practicable.
Reassess the sensitivity of the resulting linked resource rather than relying on the classifications of the source datasets.
Evaluate whether false matches or missed matches could create privacy, scientific, or participant-level consequences.
Apply access and security controls appropriate to the linked dataset's actual risk.
Review tables, figures, maps, subgroup analyses, and other outputs for disclosure risks before release.
Repeat the assessment if another dataset, linkage source, or identifying capability is later introduced.
08 · Frequently Asked Questions

Frequently Asked Questions About Privacy Risks From Data Linkage

Can two anonymous datasets become identifiable when linked?

Potentially, depending on what “anonymous” means in the applicable framework and what variables are available. Combining attributes can make records more distinctive or permit linkage with other identifying information.

What is a quasi-identifier?

It is information that may not directly identify a person by itself but can contribute to identification when combined with other attributes. Examples can include detailed age, geography, occupation, dates, or demographic characteristics.

Does removing names after linkage solve the privacy problem?

It can reduce risk but does not automatically eliminate it. The remaining linked attributes may still make individuals distinguishable, and researchers who performed the linkage may already have handled identifying information.

Can linkage reveal sensitive information that neither dataset contained explicitly?

Yes. Relationships among variables from different sources can support new inferences about health, behavior, socioeconomic circumstances, family relationships, or other characteristics.

Who should keep the linkage key?

Where feasible, access should be limited to an authorized party that genuinely needs the key for linkage or approved follow-up. Analytical researchers do not automatically need access merely because they are studying the resulting linked data.

Can publication itself create a privacy breach?

Yes. Small cells, detailed subgroup descriptions, rare combinations, maps, and other outputs can sometimes disclose individuals or small groups even when the underlying dataset is securely stored.

09 · The Bottom Line

Linkage Can Change the Privacy Risk More Than the File Size

The Bottom Line

Two datasets that are reasonably protected separately can create substantially greater privacy risks when linked because their combination may increase identifiability, reveal sensitive relationships, and enable new inferences.

Assess the linked resource as something new. Reconsider who can identify participants, what additional information becomes visible, what variables are genuinely necessary, and whether access, security, and disclosure controls remain adequate after the connection is made.

10 · Sources and Further Reading

Authoritative Guidance on Privacy and Linked Research Data

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes