03 · What You Need to Know
Demographic Variables Are Often Indirect Identifiers
A Demographic Variable Does Not Need to Name Someone to Help Identify Them
Demographics commonly function as indirect identifiers. They narrow the possible people rather than explicitly stating identity.
The ICO's guidance on indirect identification gives combinations such as age, occupation, and place of residence as examples of information that can permit a person to be identified when considered together with other information.
Current ICO anonymisation guidance similarly treats singling out and linkability as key indicators of identifiability and emphasises the importance of the richness and context of the information.
This is why indirect identifiers deserve attention even when every direct identifier has been removed.
Precision Changes Identification Risk
"Age 40–49" reveals less precise information than "age 47." "Northern region" is broader than a small municipality. "Healthcare professional" is broader than "the only pediatric neurosurgeon at Hospital X."
Greater precision can be scientifically necessary, but it also reduces the number of people who match a value.
| Less Precise |
More Precise |
Why Precision Can Matter |
| Age 40–49 |
Age 47 |
Exact age narrows the possible people |
| Region |
Small municipality or postcode |
Fine geography reduces the relevant population |
| Academic |
Professor of a rare specialty |
Detailed occupation can become distinctive |
| Senior employee |
Vice president for a specific portfolio |
A unique organisational role may effectively identify one person |
| More than 10 years' experience |
27 years' experience |
Exact tenure can support matching with public profiles |
The correct response is not always to collect broad categories. If exact age is analytically necessary, collecting only age bands may damage the research. The researcher should instead justify the precision and manage the resulting identification risk.
Rarity Can Make One Demographic Attribute Highly Revealing
Some characteristics are common in one population and unusual in another.
Knowing that a participant is a teacher may reveal little. Knowing that the participant is the only Indigenous school principal in a small district may narrow the possibilities considerably.
Likewise, a rare occupation, nationality, disability, educational background, family structure, or institutional role can become identifying in a bounded population.
The identifying power of a demographic category therefore depends on its frequency in the relevant population rather than on the variable label alone.
Several Common Characteristics Can Form a Rare Intersection
The more frequent problem is not one extraordinary variable but an extraordinary combination of ordinary ones.
Consider:
- female;
- age 55–59;
- professor;
- engineering;
- university in Municipality X;
- more than 25 years of service.
Each category may contain multiple people. Their intersection may contain one.
This is sometimes discussed in terms of quasi-identifiers: variables that can be used together to distinguish records or link them to identified information.
The broader principle is straightforward. Do not ask only whether each demographic field is identifying. Ask whether the combination creates a recognisable person.
The Same Demographic Information Can Be Safe in One Dataset and Revealing in Another
Context changes the denominator.
An exact age may have little identifying value in a national dataset containing millions of people. It can become revealing in a small sample drawn from a bounded population.
Current ICO guidance makes the same contextual point when explaining singling out: information such as year of birth may distinguish someone in one group but not another.
There is therefore no permanent classification such as "age is safe" or "occupation is identifying." Risk arises from information plus context.
Geography Is Particularly Powerful Because It Connects Data to Populations
Location variables can sharply reduce the number of possible people. Country may reveal little. Province may reveal more. Municipality, postcode, neighbourhood, workplace, or exact coordinates may narrow the population dramatically.
Geography can also facilitate linkage with electoral rolls, professional registers, institutional directories, property records, social-media profiles, or other location-based information where those sources are available.
Researchers should therefore collect geographic precision according to analytical need rather than by default.
Dates Can Behave Like Demographics
Dates are not always thought of as demographic information, but they frequently interact with demographics in identification.
Year of appointment, graduation year, migration year, admission date, date of an incident, or exact birth date can connect a research record to public or administrative information.
HHS HIPAA de-identification guidance illustrates the identification potential of detailed dates and geography by including most elements of dates directly related to an individual and detailed geographic subdivisions among the identifiers addressed by its Safe Harbor method. That standard applies specifically to protected health information under HIPAA and should not be treated as a universal research list.
The useful lesson is that temporal and geographic precision can materially strengthen a linkage attack.
Public Professional Profiles Can Make Workplace Demographics Easy to Link
Academic and professional research populations can be particularly linkable because substantial demographic and career information is already public.
Institutional websites may list names, ranks, departments, qualifications, research areas, leadership positions, and appointment histories. Professional profiles can add education and employment dates.
A research dataset containing the same variables may therefore be much easier to link than researchers expect.
This is an example of re-identification risk created by combining information.
Collecting Demographics and Reporting Demographics Are Separate Decisions
A researcher may legitimately need detailed demographics for analysis but not need to publish every variable at full precision.
For example, exact age might be used as a continuous covariate during analysis while participant characteristics are reported using broader age categories. Detailed geographic information might be required for modelling but omitted or generalised in a public dataset.
Researchers should therefore distinguish:
Analytical dataset
Contains the level of demographic detail legitimately required to conduct the approved analysis, subject to appropriate access controls.
Released or published output
Contains the demographic detail needed to communicate and support the findings without unnecessarily increasing identification risk.
This distinction allows researchers to preserve valid analysis without assuming that every collected variable must appear in every public table.
Demographic Tables Can Identify People Through Small Cells
Participant-characteristics tables often cross several demographic dimensions. This can create categories containing one or very few participants.
Imagine a study in which Table 1 shows:
- one participant aged over 65;
- one participant from Institution B;
- one participant holding the rank of dean; and
- one participant with more than 30 years' experience.
If all four descriptions refer to the same person, readers may be able to reconstruct that participant's profile even if the table never explicitly connects the rows.
Researchers should review demographic reporting cumulatively rather than assuming that separate columns prevent inference.
Demographics Can Identify Participants Through Quotations
Qualitative researchers often attribute quotations using descriptors such as:
"Female, 42, senior lecturer, public university."
These labels help readers interpret the quotation. They can also function as a compact identification key.
Where the population is small, researchers should ask whether every descriptor is necessary for interpreting that particular quotation. Attribution can sometimes use broader categories or omit a variable that adds identification risk without adding analytical meaning.
Intersectional Analysis Creates a Real Methodological Tension
Researchers may need combinations of demographic variables precisely because social experiences differ at their intersections. Collapsing categories too aggressively can erase meaningful differences and obscure underrepresented groups.
Privacy protection therefore cannot be reduced to "report fewer demographics."
The task is to preserve analytically meaningful distinctions while managing disclosure risk. In some cases this may require controlled-access data, careful aggregation, suppression of particular public cells, qualitative contextualisation without exact descriptors, or explicit discussion of why certain subgroup analyses cannot safely be reported.
This is a genuine trade-off. Privacy protection that makes marginalised groups statistically invisible is not methodologically neutral.
Demographic Information Can Identify Third Parties Too
A participant may describe someone else as "my 71-year-old husband, the former mayor of Municipality X." Even if the research participant remains anonymous, the third party may be readily identifiable.
Qualitative and mixed-methods researchers should therefore review demographic and contextual information concerning nonparticipants as well as enrolled participants.
Do Not Collect Demographics Simply Because They Are Standard
Questionnaires often include age, gender, marital status, education, income, occupation, location, and other characteristics by habit.
Each variable should have a purpose. If a demographic characteristic will not contribute to sampling, analysis, interpretation, confounding control, equity assessment, or another legitimate study function, researchers should reconsider whether it needs to be collected.
This follows the broader principle of minimising unnecessary identifying information.
Watch Out
Do not solve demographic disclosure risk by automatically replacing every detailed variable with broad categories before considering the research question. Collect and retain the precision the study genuinely needs, then manage access and public reporting according to the identification risk of each use.