01 · The Question
How Do You Know Whether You Are Collecting Too Much Identifying Information?
A registration form asks for a participant's full name, email address, telephone number, exact age, department, job title, and location. Each field seems potentially useful. Taken together, however, they create a detailed identity profile.
Researchers often focus on protecting identifiers after collection. An earlier question may be even more useful: did the study need to collect each identifier in the first place?
The appropriate amount is not always zero. Follow-up studies, longitudinal designs, incentive distribution, clinical procedures, record linkage, withdrawal mechanisms, and other legitimate research functions may require identifying information. The task is to distinguish what the study needs from what is merely convenient to have.
03 · What You Need to Know
Data Minimisation Starts Before the First Participant Enrols
Begin With Purpose, Not With a Standard Demographic Form
The easiest way to collect unnecessary personal information is to begin with last year's questionnaire and leave every field in place.
A better approach starts with purpose. For every item that identifies or helps identify a participant, ask what function it serves in this particular study.
A name might be needed for consent documentation or participant communication. An email address might be necessary for follow-up. A telephone number might be unnecessary if all communication occurs by email. Exact date of birth might be excessive if the analysis needs only broad age categories.
Under the UK GDPR's data-minimisation principle, for example, personal data should be adequate, relevant, and limited to what is necessary for the purposes for which they are processed. ICO guidance translates this into a practical instruction: identify the minimum amount of personal data needed to fulfil the purpose and hold that much, but no more.
That is a useful design question even where UK law does not apply. The applicable legal obligations will vary by jurisdiction, but unnecessary identifiers generally create privacy and governance burdens without improving the science.
"Necessary" Does Not Mean "Might Be Useful Someday"
Research produces uncertainty, and researchers understandably want flexibility. That does not make every potentially useful variable necessary.
The ICO's research guidance makes a useful distinction: necessary processing need not be absolutely indispensable, but it should be a targeted and proportionate way of achieving the research purpose. If the same purpose can reasonably be achieved through less intrusive means, collecting additional personal data requires stronger justification.
For example, collecting participants' personal mobile numbers because researchers might decide to conduct follow-up interviews later is weaker justification than collecting them because follow-up interviews are an approved component of the protocol.
The difference is between a defined purpose and speculative convenience.
Collecting No Identifier at All May Sometimes Be Possible
If a study does not require follow-up, record linkage, verification, incentives tied to identity, or another participant-specific function, researchers should consider whether identifying information is needed at all.
ICO research guidance explicitly advises researchers first to consider whether research can be conducted without personal data and, where possible, using anonymous information.
This does not mean every study should be anonymous. Many cannot be. It means researchers should not assume that identity is a default variable.
The distinction between anonymous and confidential research becomes important here. If identity is genuinely unnecessary, designing the study so the researcher never receives it may eliminate an entire category of confidentiality risk.
Collect the Least Identifying Version That Answers the Research Question
Data minimisation is not only about whether a variable is collected. It also concerns precision.
Suppose age matters analytically. Do you need:
- date of birth;
- exact age;
- five-year age bands; or
- broad age categories?
The scientifically appropriate answer depends on the analysis. A study examining age-related developmental change may legitimately require greater precision than a study using age only to describe the sample.
The same question applies to geography. An exact home address may be essential for environmental exposure modelling but excessive for a study that needs only region of residence.
Data minimisation therefore does not mean mechanically choosing the least detailed variable. The data must remain adequate for the research purpose. The objective is the least identifying level of detail that still permits the study to do what it legitimately needs to do.
Distinguish Research Variables From Administrative Identifiers
Some identifying information is analytically important. Other information exists only to operate the study.
Research information
Information needed to answer the research question, test hypotheses, describe the sample, control confounding, conduct planned analyses, or otherwise achieve the scientific purpose.
Administrative information
Information needed for recruitment, scheduling, follow-up, incentives, consent administration, withdrawal requests, or other study operations but not necessarily for analysis.
This distinction matters because administrative identifiers often do not need to travel with the research dataset. If email addresses are needed only to schedule interviews, there may be little reason for analysts to receive them alongside interview responses.
Researchers can therefore ask not only whether an identifier must be collected, but also whether it must be stored with the substantive research data.
Every Additional Identifier Expands the Identification Surface
An identifier does not have to be a name to increase identifiability. Exact age, occupation, geographic location, institutional role, event dates, and demographic details may narrow the possible participants.
Collecting several such variables can create distinctive combinations even when none identifies anyone alone. This is why researchers should consider direct and indirect identifiers together.
A useful question is therefore not merely, "Is this field sensitive?" It is also, "What does this field add to the possibility of identifying someone when combined with everything else I collect?"
Small Samples Need Particular Restraint
Detailed demographics may be scientifically useful, but their identifying power changes with population size.
Recording exact age, specialised occupation, department, seniority, gender, and institution may be unremarkable in a national survey involving thousands of participants. The same variables in a study of 12 employees from one organisation may describe individuals almost by name.
Researchers planning small or distinctive samples should therefore examine whether every demographic detail is necessary and whether broader categories would still answer the research question.
This does not justify altering scientifically necessary variables simply to make a dataset look safer. It means identification risk should form part of the study-design decision rather than being discovered at publication.
Collecting an Identifier Creates Obligations Beyond Secure Storage
Once identifying information enters the research system, researchers may need to consider lawful processing, transparency, access control, retention, disclosure, participant expectations, security, and eventual deletion or archival treatment under the rules applicable to their study.
Data minimisation can therefore simplify governance. Information that was never collected cannot be leaked from the research database, accidentally included in an export, unnecessarily shared with collaborators, or retained long after its purpose has ended.
That does not mean "collect nothing." Inadequate data can also compromise research validity. The UK GDPR formulation is deliberately balanced: personal data should be adequate and relevant as well as limited to what is necessary.
More Variables Can Also Create More Re-Identification Opportunities
A detailed dataset may become easier to link with information available elsewhere. Exact dates, locations, job titles, and demographic characteristics can provide matching variables for re-identification through other datasets.
Researchers planning eventual data sharing should therefore consider the downstream consequences of collection choices. It is easier to avoid collecting an unnecessary highly identifying variable than to preserve its full analytical detail later while somehow making its identifying power disappear.
There Is No Universal Maximum Number of Identifiers
Data minimisation is not a numerical quota. Collecting three unnecessary identifiers is not acceptable merely because another study collects ten, and a study requiring several identifiers is not automatically excessive.
The appropriate amount depends on purpose, proportionality, participant population, study procedures, analytical requirements, applicable law, ethics requirements, and available alternatives.
| Information |
Possible Legitimate Purpose |
Question to Ask |
| Full name |
Consent administration, participant-specific follow-up |
Does the analysis need the name, or can it remain separate? |
| Email address |
Scheduling, follow-up, incentive delivery |
Is email the chosen contact method, and how long must it be retained? |
| Telephone number |
Participant contact |
Is a second contact channel genuinely necessary? |
| Date of birth |
Precise age calculation or record linkage |
Would exact age or an age band accomplish the purpose? |
| Home address |
Geospatial exposure analysis or necessary correspondence |
Is exact location required, or would a broader geographic unit suffice? |
| Employer or department |
Sampling, stratification, organisational analysis |
Does this level of organisational detail materially serve the research question? |
| Detailed demographics |
Planned subgroup analysis or confounding control |
Are the categories analytically justified and proportionate to identification risk? |
06 · What This Means for You
Make Every Identifier Earn Its Place in the Dataset
Before finalising a questionnaire, interview form, case-report form, registration system, or data-extraction template, review every field that identifies or helps identify participants.
A simple decision framework
If the study can achieve its purpose without identifying participants
Consider an anonymous design rather than collecting identity by default.
If an identifier is required only for recruitment, scheduling, incentives, or follow-up
Consider collecting and storing it separately from substantive research responses.
If a detailed variable is analytically necessary
Retain the precision the research genuinely requires and document why it is necessary.
If a broader category answers the same research question
Consider collecting the broader information rather than unnecessarily precise personal data.
If you cannot explain why an identifying field is being collected
Do not treat "we might need it later" as sufficient justification without a defined purpose.
A useful protocol table can document each identifying variable, its purpose, level of precision, who needs access, where it is stored, and when that need ends. That small exercise tends to expose unnecessary fields remarkably quickly. Forms, like literature reviews, have a habit of accumulating material nobody remembers adding.