Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Is Re-Identification Risk, and Why Does It Increase When Datasets Are Combined?

Re-identification occurs when information is connected back to a person after identifying information has been removed or obscured. Risk can increase when datasets are combined because overlapping variables may provide enough clues to reconstruct an identity.

278
Re-Identification Risk and Data Linkage Guide 278 of 398
01 · The Question

How Can Combining Two Apparently Safe Datasets Reveal Someone's Identity?

A researcher removes names, email addresses, and participant numbers before sharing a dataset. The remaining file contains age, occupation, municipality, and several research variables. Nothing seems to identify anyone directly.

Then someone obtains another dataset containing names alongside age, occupation, and municipality. Suddenly, the variables that looked harmless in the first file can function as a bridge between the two.

This is one form of re-identification risk. Removing obvious identifiers can break the direct connection between a record and a person, but other information may allow that connection to be reconstructed.

02 · The Short Answer

Re-Identification Reconstructs a Link Between Data and a Person

In Brief

Re-identification risk is the possibility that information thought to be de-identified, anonymized, pseudonymized, or otherwise separated from identity can be connected back to a particular person.

Combining datasets can increase that risk because overlapping characteristics provide additional information for matching, singling out, or inferring identity. A variable that reveals little by itself may become highly identifying when connected with another dataset.

03 · What You Need to Know

Re-Identification Risk Comes From What the Data Can Be Connected To

Re-Identification Re-Establishes an Identity Connection

NIST defines re-identification as a process in which information is attributed to de-identified data in order to identify the individual to whom those data relate. It also describes the concept more generally as re-establishing the relationship between identifying data and a data subject.

This can happen in several ways. A code key might be recovered. An unusual combination of variables might point to one person. Records may be matched with another dataset containing names. Contextual knowledge may reveal who a record describes.

The central idea is the same: an identity connection that was removed, hidden, or unavailable becomes available again.

Removing Direct Identifiers Does Not Remove Every Matching Variable

Suppose Dataset A contains no names but includes:

  • age;
  • sex;
  • occupation;
  • municipality; and
  • year of appointment.

None may identify a person directly. They are potential indirect or quasi-identifying variables.

Now suppose Dataset B contains names together with occupation, municipality, and year of appointment. Those overlapping variables can provide a linkage route. If only one person has the same combination in both datasets, the previously nameless record may be connected to a named individual.

The identifying information was not necessarily hidden somewhere inside Dataset A. It emerged from the relationship between A and B.

Combining Data Adds Constraints to the Identity Puzzle

Imagine trying to identify one participant from a population of 50,000 people.

Knowing only that the participant is 45 years old may leave hundreds of candidates. Knowing that the participant is 45, works as a dentist, lives in a particular municipality, and received a professional award in a particular year may leave far fewer.

Each additional matching characteristic constrains the set of possible people. If another source associates those characteristics with a name, the identification problem may become much easier.

This is why a participant can be identified without their name appearing in the original research dataset.

Auxiliary Information Can Come From Many Places

The second source does not have to be another formal research dataset. Relevant additional or auxiliary information might come from:

  • administrative records;
  • professional directories;
  • institutional websites;
  • public registers;
  • social-media profiles;
  • news reports;
  • commercial datasets;
  • previously released research data;
  • data breaches; or
  • personal knowledge of the research population.

The relevant question is therefore broader than "What other dataset are we sharing?" Researchers should consider what information a realistic recipient already possesses or could reasonably obtain.

Current ICO anonymisation guidance explicitly requires consideration of whether another person could identify individuals from the information itself or by combining it with other information they possess or may obtain.

More Datasets Do Not Automatically Mean More Risk, but They Can Create More Linkage Opportunities

It would be too simplistic to say that combining any two datasets always increases identification risk. If the datasets contain no useful overlapping information, linkage may add little. Strong transformations, aggregation, controlled access, or other safeguards may also constrain identification.

The concern arises when combined information makes people more distinguishable or supplies a bridge to identity.

A dataset containing broad age groups and general research outcomes may have limited identification value. Add precise geography, employment records, event dates, or another source containing corresponding identifiers, and the situation can change materially.

Re-identification risk therefore depends on the informational relationship among datasets, not simply their number.

Rare Combinations Are Particularly Revealing

A common combination such as "female, age 30–39, teacher" may correspond to many people. A combination such as "female, age 67, university president, municipality X" may correspond to one.

This is one reason rare values and unusual combinations deserve attention during anonymization. The more distinctive a record becomes relative to the underlying population, the easier it may be to single out.

Small research populations can amplify this effect. When only a handful of people satisfy the inclusion criteria, even broad demographic information may substantially narrow the possibilities.

Linkage Can Reveal More Than Identity

Re-identification is concerning not only because someone may learn a participant's name. Once a research record is linked to an identified person, the recipient may also learn the sensitive attributes associated with that record.

Imagine Dataset A contains demographic information and a sensitive research outcome but no names. Dataset B contains names and overlapping demographic information but no sensitive outcome. Matching the records may reveal both who the participant is and the sensitive information recorded about them.

Identity disclosure A record is connected to a particular person.
Attribute disclosure Previously unknown information about a person can be inferred or learned, potentially even when their exact record is not fully reconstructed.

Privacy-risk assessments therefore need to consider what a successful or partial linkage would reveal, not merely whether a name can be attached to a row.

Pseudonymized Data Have an Intentional Re-Identification Route

Not every ability to restore identity is an attack or failure. In longitudinal research, researchers may deliberately retain a protected code key so participants can be linked across study waves.

Those data are generally better described as pseudonymized or coded rather than anonymous. The additional information is intentionally retained so authorised people can restore attribution when necessary.

The distinction between pseudonymization and anonymization matters because re-identification through an authorised code key is conceptually different from reconstructing identity from data intended to be anonymous.

Re-Identification Risk Depends on the Release Model

The same dataset can create different risks depending on where and how it is made available.

A restricted environment may limit users, prohibit external linkage, monitor outputs, and control which additional datasets are available. Public release makes the information available to an unknown audience with potentially diverse external data and computational resources.

NIST's guidance on de-identifying datasets consequently recommends considering the data-sharing model itself, including public release, protected enclaves, query interfaces, and other arrangements, rather than treating de-identification as a transformation detached from the release environment.

Watch Out

A dataset that presents an acceptable identification risk inside a controlled research environment should not automatically be assumed suitable for unrestricted public release. The recipients, available auxiliary information, linkage opportunities, and controls may be completely different.

Re-Identification Risk Changes Over Time

External information does not remain static. New datasets are released, online profiles expand, computational tools improve, and records that were once difficult to obtain may become searchable.

This means an anonymization or disclosure-risk assessment may need reconsideration when the information environment changes materially. The relevant question is not whether a dataset was judged safe once, but whether the assumptions underlying that judgment still hold for its current use and release environment.

NIST notes that research has demonstrated that some de-identified data can sometimes be re-identified, while its later guidance recommends evaluating disclosure risks and, where appropriate, conducting re-identification studies to assess those risks.

Re-Identification Is a Risk, Not an Automatic Fate of Every Dataset

The existence of published re-identification examples does not mean every de-identified dataset can inevitably be reconstructed. Risk varies with data granularity, population characteristics, available auxiliary information, access conditions, anonymization techniques, and realistic capabilities of potential recipients.

Researchers should therefore avoid both extremes: assuming that deleting names makes re-identification impossible, or assuming that anonymization is futile because some datasets have been re-identified.

The appropriate task is risk assessment.

Situation What Changes? Possible Effect on Re-Identification Risk
Names are removed A direct linkage route disappears Risk may decrease, but indirect linkage can remain
Detailed geography is added Records become more geographically distinctive Risk may increase
A second dataset shares several variables New matching opportunities appear Risk may increase substantially for distinctive matches
Data are aggregated Individual-level detail is reduced Risk may decrease, depending on group sizes and outputs
Access moves from controlled to public Recipients and auxiliary information become harder to constrain Risk may increase
A new public database becomes available Additional linkage information enters the environment Previous assumptions may need reassessment
04 · A Practical Example

How Two Datasets Can Reveal What Neither Reveals Alone

Hypothetical Example

A Study of Employee Well-Being

A research team releases a dataset containing employee age, sex, job category, work location, year hired, and well-being scores. Names, employee numbers, and email addresses have been removed.

Dataset A: research data One record describes a 52-year-old female laboratory director working at Site C who was hired in 2019. The record also contains a sensitive well-being score.
Dataset B: public staff directory An institutional website lists staff names, positions, locations, and appointment histories.
Find the overlap The job category, work location, and appointment year correspond to only one person in the directory.
Link the records The recipient can plausibly associate the nameless research record with the named employee.
Infer the sensitive attribute Once the match is made, the well-being score in Dataset A becomes associated with that employee.

The problem was not simply that Dataset A contained "too many identifiers." Its risk depended partly on what other information could be brought to it. A different release environment or less granular variables might produce a different assessment.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Re-Identification Risk

Misconception

If Names Are Removed, Re-Identification Is Impossible

Removing names blocks one identification route. Indirect identifiers, distinctive combinations, contextual information, code keys, and external datasets may still provide others.

Misconception

Only Another Research Dataset Can Create Linkage Risk

Auxiliary information can come from administrative records, public websites, professional directories, social media, news reports, commercial sources, or personal knowledge. What matters is whether overlapping information supports identification.

Misconception

Combining Any Two Datasets Automatically Identifies People

No. Linkage requires useful overlap, sufficient distinctiveness, or another identification mechanism. Combining datasets can increase risk, but the size of that increase depends on their contents and context.

Misconception

If Identity Is Not Revealed, No Privacy Harm Can Occur

Privacy risk can include attribute disclosure, inference, singling out, or learning additional information about a person. A risk assessment should consider what the linkage permits someone to learn, not only whether it produces a name.

Misconception

A Dataset De-Identified Once Is Safe for Every Future Release

Recipients, available external information, technology, and intended uses can change. A dataset prepared for restricted access may require a different assessment before public release or linkage with new sources.

06 · What This Means for You

Assess What Your Dataset Can Be Linked To, Not Just What It Contains

When preparing research data for sharing, start with the dataset itself but do not end there. Consider the surrounding information environment and the people who will receive the data.

A simple decision framework

If the dataset contains detailed indirect identifiers
Assess their uniqueness individually and in combination rather than assuming removal of names is sufficient.
If another dataset contains overlapping variables
Consider whether those variables could be used to match records or infer sensitive attributes.
If the data concern a small or distinctive population
Treat contextual knowledge and rare combinations as potentially important linkage routes.
If the data will move from restricted to wider access
Reassess re-identification risk for the new recipients and information environment.
If useful research data cannot be made sufficiently anonymous for public release
Consider controlled access, additional transformation, aggregation, synthetic data, query systems, or another sharing model appropriate to the research and governing requirements.

There is often a trade-off between preserving analytical detail and reducing disclosure risk. NIST's current guidance treats de-identification as a governance problem as well as a technical one, emphasizing both transformation techniques and decisions about how data will be shared.

07 · A Quick Checklist

Check the Linkage Environment Before Sharing Research Data

Before releasing de-identified or anonymized research data, check:
Identify variables that could support matching, including age, dates, geography, occupation, institutional role, and unusual characteristics.
Look for rare combinations that may single out individual participants.
Identify realistic public, administrative, commercial, research, or institutional sources containing overlapping information.
Consider what information intended recipients already possess or could reasonably obtain.
Assess what sensitive attributes could be inferred if records were successfully linked.
Evaluate the actual release model rather than assuming a dataset suitable for controlled access is suitable for public release.
Consider whether generalisation, suppression, aggregation, or another disclosure-control technique can reduce unnecessary linkage opportunities.
Reassess risk when new datasets, recipients, technologies, or uses materially change the information environment.
08 · Frequently Asked Questions

Frequently Asked Questions About Re-Identification

What does re-identification mean?

Re-identification generally refers to re-establishing a relationship between data from which identity has been removed or obscured and the individual to whom those data relate. The precise terminology can vary across technical and regulatory frameworks.

Why does combining datasets increase re-identification risk?

Combined datasets may provide more characteristics for matching and more information for distinguishing individuals. Overlapping variables can create a bridge between an unnamed research record and another source containing identity information.

Can public information be used to re-identify research data?

Potentially. Institutional websites, public registers, professional profiles, news reports, and other public sources may contain information that overlaps with research variables and supports linkage.

Does removing direct identifiers prevent re-identification?

It can substantially reduce risk, but it does not necessarily eliminate it. Indirect identifiers, distinctive combinations, external information, and other linkage mechanisms may remain.

Are small datasets easier to re-identify?

Not automatically, but small or distinctive populations can make combinations of characteristics more revealing because fewer people match them. The relevant issue is uniqueness and available linkage information rather than dataset size alone.

Can controlled access reduce re-identification risk?

Yes. Restricting recipients, limiting external linkage, controlling outputs, imposing contractual conditions, and using secure environments can reduce some identification opportunities. These controls do not transform identifiable information into anonymous information in every legal framework, but they can form an important part of risk management.

Does re-identification risk ever become zero?

Absolute zero risk is not the standard used by every framework. Some approaches instead assess whether identification is sufficiently unlikely or whether means of identification are reasonably likely to be used. Researchers should apply the threshold required by the governing framework.

09 · The Bottom Line

A Dataset's Identification Risk Depends Partly on the Data Around It

The Bottom Line

Re-identification risk is the possibility of reconnecting research information to individuals, and combining datasets can increase that risk when overlapping variables create new routes for matching, singling out, or inference.

Do not assess a dataset in isolation. Consider realistic auxiliary information, likely recipients, distinctive combinations, the release environment, and what a successful linkage could reveal. Data that look innocuous alone may tell a very different story once another dataset joins the meeting.

10 · Sources and Further Reading

Authoritative Sources on Re-Identification and Data Linkage

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes