Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Large-Scale Data Collection Become Ethically Problematic Even When No Individual Data Item Is Sensitive?

A dataset can become ethically sensitive even when none of its individual records appears sensitive. Aggregation, linkage, inference, scale, persistence, and re-identification can create information and risks that were absent from each data point in isolation.

362
Ethics of Large-Scale Data Collection Guide 362 of 398
01 · The Question

Can harmless pieces of information become sensitive when researchers collect enough of them?

Consider a dataset containing ordinary public facts: timestamps, general locations, likes, purchases, posts, follows, transport trips, website interactions, or attendance at public events.

Look at one record and little seems sensitive.

Now collect millions of records, connect them by person, order them through time, combine them with other datasets, and use statistical or machine-learning methods to infer patterns. The resulting dataset may reveal routines, relationships, beliefs, vulnerabilities, identities, or characteristics that no single record disclosed.

Ethical sensitivity can therefore emerge from the structure of a dataset, not merely from the sensitivity of its individual fields.

02 · The Short Answer

Yes, aggregation can create risks that individual records do not contain

In Brief

Yes. Large-scale data collection can become ethically problematic even when individual data items appear harmless, because aggregation, linkage, longitudinal tracking, inference, re-identification, and broad reuse can reveal sensitive information or create risks that are absent from isolated records.

Researchers should assess the dataset as a whole: what it allows them or others to infer, whom it can identify, how many people could be affected, whether all collected variables are necessary, how long the data persist, and what consequences could follow from breach, misuse, publication, or secondary analysis.

03 · What You Need to Know

The ethical properties of a dataset are not simply the sum of its rows

Aggregation can reveal patterns invisible in individual records

A single timestamp may reveal almost nothing. Thousands of timestamps associated with the same person can reveal routines.

One location may be innocuous. A sequence of locations can suggest where someone lives, works, worships, receives medical care, socializes, or spends nights.

One social connection may reveal little. A network of connections can expose communities, relationships, organizational structures, or affiliations.

HHS advisory guidance on internet research explicitly recognizes this problem. It notes that online data can be mined and matched, that partial identifiers can be combined, and that multiple datasets can be aggregated to produce surprising or novel information about individuals.

Sensitivity can be inferred rather than collected

Researchers sometimes describe a dataset as non-sensitive because it contains no field explicitly labeled health status, religion, political affiliation, income, or sexual orientation.

That description can be misleading if those characteristics can be inferred from other variables.

Explicit sensitive information A sensitive characteristic appears directly in the collected data.
Inferred sensitive information A sensitive characteristic is estimated or reconstructed from combinations of information that may appear non-sensitive individually.

The ethical assessment should consider both.

Scale changes the number of people exposed to risk

A low probability of harm per person can still matter when a dataset includes millions of people.

Large-scale collection increases the number of individuals potentially affected by data breaches, incorrect classifications, unauthorized secondary use, re-identification, discriminatory applications, or disclosure.

Researchers should therefore avoid assessing risk only at the level of a typical individual record. Population-level exposure also matters.

Linkage can change identifiability

A dataset may appear de-identified when viewed alone but become identifying when linked with another source.

HHS guidance defines identifiable private information in relation to whether identity is or may readily be ascertained or associated with the information, and its internet-research guidance emphasizes the distinctive ability of online data to be mined and matched.

This means researchers should ask not merely whether their dataset contains names but whether available variables permit practical linkage to identifiable external information.

Public information can still become a surveillance-like dataset

Suppose a researcher collects only information that individuals deliberately made public. At the level of each record, there may be little expectation of secrecy.

Aggregation can nevertheless transform those traces into a detailed longitudinal representation of a person's behavior.

This is a central reason the distinction between public accessibility and ethically appropriate research use matters. Research can make public information dramatically easier to search, compare, profile, and interpret than it was in its original scattered form.

More data are not automatically better data

Large datasets create a familiar temptation: collect everything now because storage is cheap and decide what matters later.

That approach can increase privacy and security risks without improving the research.

The Association of Internet Researchers' guidance encourages researchers to consider whether collected information is necessary and proportionate to the research aim and whether unnecessary data can be deleted. Data minimization therefore applies to big-data research just as it does to small studies.

A billion unnecessary variables do not become methodologically necessary through enthusiasm.

Longitudinal data can become more revealing over time

Repeated observations can create information that cross-sectional data cannot.

Individual observation What repeated collection may reveal Potential ethical concern
Approximate location Home, workplace, routines, regular destinations Identification, surveillance, sensitive-location inference
Public post Changing opinions, emotional patterns, relationships, life events Profiling and sensitive inference
Purchase category Habits, financial circumstances, possible health or lifestyle patterns Unexpected secondary inference
Social connection Communities, organizations, close relationships, network position Exposure of third parties and affiliations
Timestamp Daily routines and absence patterns Behavioral tracking and identification

The ethical assessment should therefore consider temporal depth as well as sample size.

Third parties can enter the dataset without being intentionally studied

Large-scale collection frequently captures information about people who are not the primary unit of analysis.

A person's post may mention a spouse, colleague, child, patient, employer, or friend. Social-network data inherently describe relationships between people. Photographs and videos may include bystanders.

As datasets grow, the number of indirectly represented people can also grow. Researchers should consider whether information authored by one person reveals identifiable or sensitive information about another.

Errors can scale too

Large datasets can create an aura of precision. Yet automated classifications, entity matching, demographic inference, sentiment analysis, geolocation, and identity resolution can all be wrong.

A false inference affecting one record may be a small analytical error. Applying the same flawed inference to millions of people can systematically misrepresent groups or produce biased conclusions.

Ethical assessment should therefore consider the consequences of algorithmic error, not merely privacy.

Group harms can occur without individual identification

Researchers may successfully anonymize every person and still produce findings that affect identifiable groups.

A model might associate a neighborhood, occupation, linguistic community, ethnic group, online community, or other population with stigmatized behavior. Publication can influence how outsiders treat that group even when no individual record is disclosed.

The ethics of research involving online communities therefore includes collective consequences alongside individual confidentiality.

Large datasets create larger security targets

Scale changes the consequences of a breach.

A compromised file containing twenty anonymous aggregate observations is different from a person-level dataset containing millions of records, persistent identifiers, networks, locations, and longitudinal histories.

Researchers should match security controls to the consequences of unauthorized access. This may include reducing identifiers, separating linkage keys, restricting access, encrypting sensitive files where appropriate, maintaining access logs, defining retention periods, and avoiding unnecessary local copies.

Secondary use can drift far from the original research purpose

Large datasets are expensive to build and therefore attractive for reuse. New questions emerge, collaborators request access, and technologies make previously impossible analyses feasible.

That scientific value can be substantial. It also creates the possibility of purpose expansion.

A dataset collected to study traffic patterns might later support behavioral profiling. Public posts collected for linguistic analysis might later be used to infer mental-health characteristics. Researchers should consider whether proposed secondary uses remain consistent with consent, ethics approval, licences, access conditions, and applicable law.

Web scraping can accelerate every one of these problems

Automated collection is one common route to very large datasets. Scraping can turn millions of public traces into a structured research resource rapidly.

The ethical responsibilities involved in web scraping public information therefore include considering what aggregation itself creates, not merely whether each source record was accessible.

Public availability can affect formal regulatory status without eliminating broader ethics

Under the U.S. revised Common Rule, certain secondary research using identifiable private information that is publicly available may qualify for an exemption. HHS's discussion of the revised rule recognizes publicly available archives and records as examples relevant to that exemption.

Formal regulatory treatment matters and researchers should apply it accurately.

But an exemption from a particular review requirement should not be translated into the claim that aggregation, security, profiling, group harm, or data minimization no longer matter. Those are broader research-design questions.

Watch Out

Do not assess a large dataset by opening one row and asking whether that row looks sensitive. Ask what the complete dataset makes possible.

04 · A Practical Example

Nothing in the dataset looks sensitive until the records are connected

Hypothetical Example

Studying public mobility traces

A researcher collects publicly visible check-ins associated with pseudonymous accounts. Each record contains an account identifier, location, and timestamp. No record contains a person's name, medical information, religion, employer, or home address.

Individual record One check-in at a café reveals little about the person behind the account.
Aggregation Hundreds of records from the same account reveal a recurring overnight location, a weekday destination, and several regularly visited venues.
Inference The pattern can plausibly reveal a home area and workplace, while repeated visits to particular specialized locations may expose sensitive associations.
Linkage The pseudonymous account can potentially be matched with another public profile containing a photograph and real name.
Ethical redesign The researcher determines that person-level trajectories are unnecessary for the research question, aggregates observations earlier, limits temporal precision, removes persistent identifiers where possible, and restricts access to the raw data.

No individual check-in changed. The ethical character of the information emerged from connecting them.

05 · What Researchers Often Get Wrong

Individually harmless data do not guarantee a harmless dataset

Misconception

If none of the variables is sensitive, the dataset is not sensitive

Combinations of ordinary variables can reveal sensitive characteristics, routines, relationships, or identities that no field states explicitly.

Misconception

If names are absent, people cannot be identified

Persistent identifiers, locations, timestamps, networks, distinctive behaviors, and linkage with external datasets can permit re-identification without names.

Misconception

Collecting more data can only improve the study

Additional variables can increase noise, security burden, privacy exposure, analytical flexibility, and opportunities for post hoc inference without improving the answer to the research question.

Misconception

If all source data are public, aggregation creates no new ethical issue

Aggregation can make information easier to search, link, profile, and interpret, creating capabilities that were not realistically available from isolated public records.

Misconception

Anonymizing individuals eliminates every possible harm

Research findings can stigmatize or disadvantage communities, neighborhoods, occupations, or other groups even when no individual is identifiable.

06 · What This Means for You

Evaluate what the assembled dataset can reveal, not only what you collected

When designing large-scale research, assess risk after imagining the records combined, linked, sorted through time, and analyzed with the methods you intend to use.

A simple decision framework

If individual records appear non-sensitive
Test whether combinations of records can reveal sensitive characteristics, identities, routines, relationships, or locations.
If persistent identifiers enable longitudinal analysis
Determine whether person-level tracking is necessary and protect or remove identifiers when it is not.
If datasets will be linked
Reassess identifiability and sensitivity after linkage rather than relying on the classification of each source dataset separately.
If large numbers of variables are available
Collect and retain only what the research can justify rather than maximizing the dataset by default.
If findings concern identifiable groups
Assess potential collective harms and avoid assuming that individual anonymization resolves every ethical concern.

Scale is not merely a larger sample size. It can change what researchers are capable of knowing about people.

07 · A Quick Checklist

Before building a large dataset, test what emerges from aggregation

Before large-scale collection or linkage, check:
Define which records and variables are necessary for the research question rather than collecting everything available.
Assess what sensitive characteristics can be inferred from combinations of otherwise ordinary variables.
Test whether persistent identifiers, locations, timestamps, networks, or behavioral patterns can permit re-identification.
Reassess privacy and identifiability after linking datasets rather than relying on the status of each source independently.
Consider how longitudinal collection changes what can be inferred about routines, relationships, and life events.
Identify third parties and groups who may be represented or affected even though they are not the primary unit of analysis.
Match access controls, security, retention, and data-sharing arrangements to the consequences of a breach involving the complete dataset.
Evaluate potential errors, biases, and group-level harms produced by automated classification or inference at scale.
Verify that secondary uses, sharing, linkage, and future analyses remain consistent with ethics approval, consent, licences, access conditions, and applicable law.
08 · Frequently Asked Questions

Questions about the ethics of large research datasets

Can non-sensitive public data become sensitive after aggregation?

Yes. Combining records can reveal routines, relationships, identities, locations, affiliations, or other characteristics that no individual data point exposes.

Does removing names make a large dataset anonymous?

Not necessarily. Persistent identifiers, timestamps, locations, network structures, rare characteristics, and linkage with external sources may permit re-identification.

Why does dataset size matter if the risk to each person is small?

Scale increases the number of people potentially affected and can enable analyses, linkages, and inferences that small datasets cannot. The aggregate consequences of a breach or systematic error may also become substantial.

Is collecting every available variable good research practice?

Not automatically. Data minimization can reduce privacy, security, and analytical risks. Researchers should be able to explain how collected variables contribute to the research purpose or another legitimate requirement.

Can anonymized research still harm a community?

Yes. Findings can stigmatize, stereotype, expose, or disadvantage identifiable groups even when no individual person is named or re-identifiable.

Does an ethics exemption mean large public datasets are ethically unrestricted?

No. An exemption addresses the requirements of a particular regulatory framework. Researchers may still need to address data minimization, security, profiling, linkage, group harms, platform conditions, responsible reporting, and applicable law.

Should researchers share large public datasets openly for reproducibility?

Not automatically. Researchers should assess whether redistribution increases identifiability, searchability, sensitivity, or legal and contractual risk. Reproducibility can sometimes be supported through code, derived data, controlled access, or other alternatives.

09 · The Bottom Line

Data can become sensitive through combination, not just collection

The Bottom Line

Large-scale data collection can become ethically problematic even when no individual record appears sensitive, because aggregation, linkage, longitudinal tracking, inference, and re-identification can create information and risks that do not exist at the level of a single data point.

Evaluate the capabilities of the completed dataset rather than judging each variable in isolation. Ask what researchers or future users can infer, who can be identified, which groups may be affected, what happens if the dataset leaks, and whether every collected element is genuinely necessary for the research.

10 · Sources and Further Reading

Authoritative guidance on aggregation, identifiability, and large-scale research data

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes