Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Personal Data Do Researchers Actually Need to Collect?

Researchers should collect personal data because a defined research or operational purpose requires them, not simply because the information might be useful. Each variable should have a defensible reason for being collected.

307
What Personal Data Should Researchers Collect? Guide 307 of 398
01 · The Question

How Much Participant Information Do You Really Need?

Age, sex, gender, email address, occupation, income, exact location, student number, medical history, ethnicity, religion: once researchers start designing a questionnaire or database, the list of potentially useful variables can grow surprisingly quickly.

The temptation is understandable. If you are already collecting data, why not ask a few additional questions in case they become useful later?

That is precisely where researchers should become more selective. Personal data should be connected to a defined research, administrative, safety, or other legitimate purpose. A variable being interesting, conventional, or easy to collect does not by itself make it necessary.

02 · The Short Answer

Collect What the Study Needs, Not Everything You Could Use

In Brief

Researchers should collect only the personal data that can be justified by a specific purpose in the study or its legitimate operation, while still collecting enough information to answer the research question properly.

The appropriate dataset depends on the research design. Before collecting a personal-data variable, ask what purpose it serves, whether it is actually needed for that purpose, and whether the same objective could reasonably be achieved with less identifying or less sensitive information.

03 · What You Need to Know

Start With Purpose, Then Decide What Data Follow From It

There is no universal list of personal data every study should collect

A longitudinal clinical study, an anonymous classroom survey, an interview study, and a population database analysis have very different information requirements. Asking which personal data researchers are allowed to collect without first asking why the data are needed reverses the decision process.

A better sequence begins with the research question and study design. Determine what needs to be measured, what information is necessary to recruit or communicate with participants, what information is required for safety or follow-up, and what records may be necessary for legitimate administrative or regulatory purposes. Only then should you decide which personal-data fields are needed.

This also means distinguishing the information needed to conduct the study from the information needed to analyze it. A participant's email address might be necessary for scheduling an interview but completely unnecessary in the analytical dataset. Contact details can often be stored separately rather than accompanying the research responses throughout the project.

Personal data are not limited to names and contact details

What legally constitutes personal data depends on the applicable jurisdiction, but researchers should not assume that removing names automatically removes all personal information. Information can identify a person directly or, in some circumstances, indirectly when combined with other information.

A research dataset might therefore contain personal data even when there is no column labeled "Name." Exact dates, detailed geographic information, institutional identifiers, online identifiers, unusual occupations, combinations of demographic characteristics, images, audio recordings, or sufficiently distinctive free-text responses can potentially contribute to identification.

Whether a particular dataset legally constitutes personal data requires assessment under the applicable framework. The practical lesson is simpler: think about identifiability across the dataset, not merely whether obvious identifiers have been removed.

Give every personal-data variable a job

One useful discipline is to require a purpose for each personal-data field before it enters the data collection instrument. That purpose might be analytical, operational, safety-related, regulatory, or methodological.

Possible variable A defensible reason might be Question to ask before collecting it
Age Eligibility depends on age, age is an explanatory variable, or age adjustment is specified in the analysis. Do you need exact age or date of birth, or would an age range be sufficient?
Email address Participants must receive appointments, follow-up surveys, study results, or compensation information. Does the email address need to remain attached to the research dataset?
Exact home address The research genuinely requires household-level location or an intervention must be delivered there. Would city, municipality, district, postal area, or another less precise location answer the research question?
Occupation Occupation is part of the hypothesis, sampling strategy, exposure assessment, or planned subgroup analysis. Do you need an exact job title, or would a broader occupational category work?
Income Socioeconomic position is substantively relevant to the research question or a specified confounder. Is exact income necessary, or would meaningful income bands provide sufficient analytical precision?
Student or employee number Records must be linked across an authorized source or participants must be tracked longitudinally. Can a study-generated identifier accomplish the task instead?

The examples are not rules about what researchers may or may not collect. Their purpose is to expose a more useful question: what level of detail is necessary for the particular function?

Demographic questions are not automatically necessary

Researchers frequently add standard demographic sections almost by reflex. Age, sex, gender, marital status, educational attainment, employment, income, ethnicity, religion, and location may all be important in particular studies. None is automatically necessary simply because demographic tables are common in published papers.

Ask what you intend to do with each variable. Will it describe the sample in a substantively meaningful way? Is it part of a hypothesis? Will it be used in sampling, adjustment, stratification, effect-modification analysis, or interpretation? Is it needed to assess representativeness or inequity? Is there another legitimate requirement?

If the only explanation is "we usually include it," reconsider the variable. Established disciplinary practice can inform study design, but habit alone does not demonstrate necessity.

Collecting enough data is also part of good research

Minimizing personal data does not mean indiscriminately deleting variables until privacy risk approaches zero. Under the GDPR formulation, for example, personal data should be "adequate, relevant and limited to what is necessary" for the purpose. Adeacy matters alongside limitation.

If age is a genuine confounder, refusing to collect any age information in the name of privacy could weaken the study. If longitudinal follow-up is essential, some mechanism for reconnecting records may be necessary. If a study examines disparities between groups, relevant demographic characteristics may be scientifically indispensable.

The goal is therefore not the smallest imaginable dataset. It is the smallest dataset that still properly fulfils the defined purpose.

Ask whether you need the exact value

Sometimes the variable is necessary but its precision is not. A researcher may need participants' ages without needing full dates of birth. Geographic context may matter without requiring exact residential addresses. Socioeconomic status may be analyzable using categories rather than exact financial figures.

This creates an important distinction between needing information about a characteristic and needing its most precise possible value. Reducing precision can sometimes preserve analytical utility while reducing identifiability or sensitivity.

Ask whether you need identifying information in the analytical dataset

A project may legitimately require identifiers at one stage without needing them everywhere. Recruitment staff may need names and contact information. Analysts may need only study IDs and research variables.

Separating contact or identifying information from analytical data can reduce unnecessary exposure. Pseudonymization may also be appropriate in some projects, although pseudonymized information generally remains personal data under GDPR-style frameworks when reidentification remains possible.

This is one reason to think about responsibility for personal research data at the design stage rather than after collection has begun. Different members of the team do not necessarily need access to the same information.

"It might be useful later" is usually not enough

Exploratory research can legitimately require flexibility, and not every future analysis can be predicted with perfect precision. That does not make unlimited collection defensible.

UK Information Commissioner's Office guidance on data minimisation explicitly advises against collecting personal data merely on the off-chance that they might become useful. Its research guidance also explains that necessity must amount to more than something being useful or habitual: the processing should be a targeted and proportionate means of achieving the research purpose.

The Philippine Data Privacy Act and its implementing rules use the related principle of proportionality. Personal data processing should be adequate, relevant, suitable, necessary, and not excessive in relation to a declared and specified purpose. The implementing rules further state that only personal data necessary and compatible with that purpose should be collected.

Different legal frameworks use somewhat different terminology, but both illustrate the same practical discipline for research design: collect information because you can explain why you need it.

Sensitive information deserves an even stronger justification

Health information, genetic or biometric information, racial or ethnic origin, political or religious information, sexual-life information, and other legally protected categories vary across jurisdictions. Their processing may trigger additional requirements.

Even apart from legal classification, some information can create greater harm if disclosed or misused. Researchers should therefore ask whether a less sensitive variable could answer the same question, whether a broader category would suffice, and whether the information needs to remain identifiable.

Watch Out

Do not assume that participant consent makes unnecessary data collection harmless. Consent and necessity answer different questions. Applicable law, ethics requirements, and institutional policies may still limit what should be collected and how it may be processed.

Data requirements can change during the research lifecycle

A field may be necessary during recruitment but unnecessary after eligibility is confirmed. Contact information may be needed during follow-up but not after the final participant communication. A linkage key may be needed until datasets are combined and then become unnecessary for most members of the research team.

For that reason, deciding what to collect is only the first step. Researchers should also determine who needs each type of information, for how long, and at which stage of the project. These questions lead directly to data minimization across the research lifecycle.

04 · A Practical Example

Turning a 12-Field Participant Profile Into What the Study Actually Needs

Hypothetical Example

A study of university students' study habits and academic stress

A researcher drafts a survey asking for full name, university email, student number, exact date of birth, age, sex, gender, home address, degree program, year level, household income, and religion before the main measures of study habits and academic stress.

Define the analysis The research questions require age group, degree program, year level, study-habit measures, and academic-stress scores. No hypotheses or planned analyses involve religion, exact residential location, household income, sex, or gender.
Separate operational needs Email addresses are needed only because participants who request a summary of findings will receive one later. They are not required for analysis.
Reduce unnecessary precision The analysis requires age categories rather than exact birth dates. The survey therefore asks for the required age information without collecting a full date of birth.
Remove unexplained variables The researcher removes fields that have no defined methodological or operational purpose rather than keeping them for unspecified future analyses.
Separate identifiers Email addresses for participants requesting results are stored separately from survey responses using an appropriate study process.

This does not mean the removed variables are inherently inappropriate research variables. A study of gender differences in academic stress could have a clear reason to collect gender. A socioeconomic inequality study could require an income measure. The justification changes because the research purpose changes.

05 · What Researchers Often Get Wrong

Common Reasons Researchers Collect More Than They Need

Misconception

Should Every Survey Have a Full Demographic Profile?

No. Demographic variables should serve a methodological, interpretive, administrative, ethical, or other legitimate purpose. A standard demographic block copied from previous studies can quietly accumulate information that the new project never uses.

Misconception

If a Variable Might Be Useful for a Future Paper, Should You Collect It Now?

Possibly, but a merely speculative future use is a weak justification. If secondary analyses are genuinely part of the research purpose, define them sufficiently to justify the relevant information and ensure that the applicable ethical, legal, and transparency requirements are addressed.

Misconception

Does Removing Names Make Every Other Question Safe to Collect?

No. People can sometimes be identifiable through combinations of variables, and information can remain sensitive even without a name attached. Evaluate the dataset as a whole and consider whether the level of detail creates unnecessary identification or disclosure risk.

Misconception

Is More Data Always Better Science?

No. Additional variables can sometimes improve analysis, but indiscriminate collection can add participant burden, complicate governance, increase exposure in a breach, and create analytical opportunities that were never methodologically justified. Good measurement is selective, not merely abundant.

Misconception

Does Data Minimization Mean Researchers Should Collect as Little as Possible?

No. The aim is not arbitrary scarcity. The dataset must remain adequate for its purpose. A necessary confounder, eligibility variable, linkage field, or safety measure should not be omitted merely to reduce the number of columns.

06 · What This Means for You

Make Every Personal-Data Field Defend Its Place

Before finalizing a questionnaire, interview protocol, extraction form, database, or data request, review the personal-data variables one by one. The useful question is not simply, "Could this be useful?" but "What defined purpose requires this information at this level of detail?"

A simple decision framework

If the variable directly measures something required by the research question
Collect the level of information necessary for valid measurement, subject to applicable ethical and legal requirements.
If the variable is needed for eligibility, sampling, confounding control, stratification, linkage, follow-up, or safety
Document that purpose and ask whether a less identifying or less precise version would still work.
If the variable is included only because similar studies collect it
Reassess it. Convention can inform your reasoning but should not replace it.
If you cannot explain what you will do with the information
Do not collect it merely in case it becomes interesting later.

Then repeat the exercise for precision. You may need age but not date of birth, region but not street address, an occupational category but not an employer's name. The appropriate choice depends on what the research actually needs.

Finally, separate collection from access. Even where the project legitimately collects identifiable information, that does not mean every collaborator needs access to participant-level identifiers. What the project needs and what each individual team member needs are different questions.

07 · A Quick Checklist

Before Adding a Personal-Data Field

For each personal-data variable, check:
State the specific research, operational, safety, regulatory, or methodological purpose the variable serves.
Confirm that the variable is actually relevant to that purpose rather than merely potentially interesting.
Ask whether the purpose can reasonably be achieved without collecting personal data at all.
Ask whether a broader category, less precise value, or study-generated identifier would be sufficient.
Verify whether the information is legally sensitive or otherwise presents heightened privacy or participant risk.
Determine whether identifiers can be stored separately from the analytical dataset.
Define who actually needs access to the variable and during which stage of the project.
Check applicable ethics requirements, data protection law, and institutional policies before collection begins.
08 · Frequently Asked Questions

Common Questions About Collecting Personal Data for Research

Should I collect participants' names?

Only when the study has a reason to identify participants, such as scheduling, longitudinal follow-up, record linkage, compensation, or another legitimate requirement. If names are needed operationally but not analytically, consider keeping them separate from research responses.

Do I need to ask participants their age?

That depends on the study. Age may be necessary for eligibility, description, analysis, adjustment, or interpretation. If exact age or date of birth is unnecessary, a less precise measure may be sufficient.

Should every research survey collect sex and gender?

No. Collect them when they are relevant to the research or another defined purpose. Also determine which concept the study actually requires rather than treating sex and gender as interchangeable variables.

Can I collect extra variables for possible future research?

Do not assume that an unspecified possible future use is sufficient. Secondary or future research can sometimes be planned legitimately, but the collection and subsequent use must satisfy the applicable ethical, legal, and institutional requirements.

Is a coded participant ID still personal data?

It can be. Under GDPR-style frameworks, pseudonymized information remains personal data when it can be attributed to an individual using additional information. The exact legal assessment depends on the applicable framework and circumstances.

What if I discover after data collection that I did not need a variable?

Review whether there is any continuing justification for retaining or processing it and follow the applicable retention, deletion, ethics, and institutional requirements. Data minimization is an ongoing responsibility, not merely a questionnaire-design exercise.

09 · The Bottom Line

Collect Personal Data With a Reason, Not Just a Possibility

The Bottom Line

Researchers should collect enough personal data to conduct the study properly, but every personal-data variable should have a defensible connection to a defined research or legitimate operational purpose.

Start with the purpose, then determine the information and precision that purpose requires. If you cannot explain why you need a personal-data field, that is a good reason to reconsider collecting it.

10 · Sources and Further Reading

Authoritative Sources on Necessary Personal Data

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes