01 · The Question
Can harmless pieces of information become sensitive when researchers collect enough of them?
Consider a dataset containing ordinary public facts: timestamps, general locations, likes, purchases, posts, follows, transport trips, website interactions, or attendance at public events.
Look at one record and little seems sensitive.
Now collect millions of records, connect them by person, order them through time, combine them with other datasets, and use statistical or machine-learning methods to infer patterns. The resulting dataset may reveal routines, relationships, beliefs, vulnerabilities, identities, or characteristics that no single record disclosed.
Ethical sensitivity can therefore emerge from the structure of a dataset, not merely from the sensitivity of its individual fields.
03 · What You Need to Know
The ethical properties of a dataset are not simply the sum of its rows
Aggregation can reveal patterns invisible in individual records
A single timestamp may reveal almost nothing. Thousands of timestamps associated with the same person can reveal routines.
One location may be innocuous. A sequence of locations can suggest where someone lives, works, worships, receives medical care, socializes, or spends nights.
One social connection may reveal little. A network of connections can expose communities, relationships, organizational structures, or affiliations.
HHS advisory guidance on internet research explicitly recognizes this problem. It notes that online data can be mined and matched, that partial identifiers can be combined, and that multiple datasets can be aggregated to produce surprising or novel information about individuals.
Sensitivity can be inferred rather than collected
Researchers sometimes describe a dataset as non-sensitive because it contains no field explicitly labeled health status, religion, political affiliation, income, or sexual orientation.
That description can be misleading if those characteristics can be inferred from other variables.
Explicit sensitive information
A sensitive characteristic appears directly in the collected data.
Inferred sensitive information
A sensitive characteristic is estimated or reconstructed from combinations of information that may appear non-sensitive individually.
The ethical assessment should consider both.
Scale changes the number of people exposed to risk
A low probability of harm per person can still matter when a dataset includes millions of people.
Large-scale collection increases the number of individuals potentially affected by data breaches, incorrect classifications, unauthorized secondary use, re-identification, discriminatory applications, or disclosure.
Researchers should therefore avoid assessing risk only at the level of a typical individual record. Population-level exposure also matters.
Linkage can change identifiability
A dataset may appear de-identified when viewed alone but become identifying when linked with another source.
HHS guidance defines identifiable private information in relation to whether identity is or may readily be ascertained or associated with the information, and its internet-research guidance emphasizes the distinctive ability of online data to be mined and matched.
This means researchers should ask not merely whether their dataset contains names but whether available variables permit practical linkage to identifiable external information.
Public information can still become a surveillance-like dataset
Suppose a researcher collects only information that individuals deliberately made public. At the level of each record, there may be little expectation of secrecy.
Aggregation can nevertheless transform those traces into a detailed longitudinal representation of a person's behavior.
This is a central reason the distinction between public accessibility and ethically appropriate research use matters. Research can make public information dramatically easier to search, compare, profile, and interpret than it was in its original scattered form.
More data are not automatically better data
Large datasets create a familiar temptation: collect everything now because storage is cheap and decide what matters later.
That approach can increase privacy and security risks without improving the research.
The Association of Internet Researchers' guidance encourages researchers to consider whether collected information is necessary and proportionate to the research aim and whether unnecessary data can be deleted. Data minimization therefore applies to big-data research just as it does to small studies.
A billion unnecessary variables do not become methodologically necessary through enthusiasm.
Longitudinal data can become more revealing over time
Repeated observations can create information that cross-sectional data cannot.
| Individual observation |
What repeated collection may reveal |
Potential ethical concern |
| Approximate location |
Home, workplace, routines, regular destinations |
Identification, surveillance, sensitive-location inference |
| Public post |
Changing opinions, emotional patterns, relationships, life events |
Profiling and sensitive inference |
| Purchase category |
Habits, financial circumstances, possible health or lifestyle patterns |
Unexpected secondary inference |
| Social connection |
Communities, organizations, close relationships, network position |
Exposure of third parties and affiliations |
| Timestamp |
Daily routines and absence patterns |
Behavioral tracking and identification |
The ethical assessment should therefore consider temporal depth as well as sample size.
Third parties can enter the dataset without being intentionally studied
Large-scale collection frequently captures information about people who are not the primary unit of analysis.
A person's post may mention a spouse, colleague, child, patient, employer, or friend. Social-network data inherently describe relationships between people. Photographs and videos may include bystanders.
As datasets grow, the number of indirectly represented people can also grow. Researchers should consider whether information authored by one person reveals identifiable or sensitive information about another.
Errors can scale too
Large datasets can create an aura of precision. Yet automated classifications, entity matching, demographic inference, sentiment analysis, geolocation, and identity resolution can all be wrong.
A false inference affecting one record may be a small analytical error. Applying the same flawed inference to millions of people can systematically misrepresent groups or produce biased conclusions.
Ethical assessment should therefore consider the consequences of algorithmic error, not merely privacy.
Group harms can occur without individual identification
Researchers may successfully anonymize every person and still produce findings that affect identifiable groups.
A model might associate a neighborhood, occupation, linguistic community, ethnic group, online community, or other population with stigmatized behavior. Publication can influence how outsiders treat that group even when no individual record is disclosed.
The ethics of research involving online communities therefore includes collective consequences alongside individual confidentiality.
Large datasets create larger security targets
Scale changes the consequences of a breach.
A compromised file containing twenty anonymous aggregate observations is different from a person-level dataset containing millions of records, persistent identifiers, networks, locations, and longitudinal histories.
Researchers should match security controls to the consequences of unauthorized access. This may include reducing identifiers, separating linkage keys, restricting access, encrypting sensitive files where appropriate, maintaining access logs, defining retention periods, and avoiding unnecessary local copies.
Secondary use can drift far from the original research purpose
Large datasets are expensive to build and therefore attractive for reuse. New questions emerge, collaborators request access, and technologies make previously impossible analyses feasible.
That scientific value can be substantial. It also creates the possibility of purpose expansion.
A dataset collected to study traffic patterns might later support behavioral profiling. Public posts collected for linguistic analysis might later be used to infer mental-health characteristics. Researchers should consider whether proposed secondary uses remain consistent with consent, ethics approval, licences, access conditions, and applicable law.
Web scraping can accelerate every one of these problems
Automated collection is one common route to very large datasets. Scraping can turn millions of public traces into a structured research resource rapidly.
The ethical responsibilities involved in web scraping public information therefore include considering what aggregation itself creates, not merely whether each source record was accessible.
Public availability can affect formal regulatory status without eliminating broader ethics
Under the U.S. revised Common Rule, certain secondary research using identifiable private information that is publicly available may qualify for an exemption. HHS's discussion of the revised rule recognizes publicly available archives and records as examples relevant to that exemption.
Formal regulatory treatment matters and researchers should apply it accurately.
But an exemption from a particular review requirement should not be translated into the claim that aggregation, security, profiling, group harm, or data minimization no longer matter. Those are broader research-design questions.
Watch Out
Do not assess a large dataset by opening one row and asking whether that row looks sensitive. Ask what the complete dataset makes possible.
07 · A Quick Checklist
Before building a large dataset, test what emerges from aggregation
Before large-scale collection or linkage, check:
Define which records and variables are necessary for the research question rather than collecting everything available.
Assess what sensitive characteristics can be inferred from combinations of otherwise ordinary variables.
Test whether persistent identifiers, locations, timestamps, networks, or behavioral patterns can permit re-identification.
Reassess privacy and identifiability after linking datasets rather than relying on the status of each source independently.
Consider how longitudinal collection changes what can be inferred about routines, relationships, and life events.
Identify third parties and groups who may be represented or affected even though they are not the primary unit of analysis.
Match access controls, security, retention, and data-sharing arrangements to the consequences of a breach involving the complete dataset.
Evaluate potential errors, biases, and group-level harms produced by automated classification or inference at scale.
Verify that secondary uses, sharing, linkage, and future analyses remain consistent with ethics approval, consent, licences, access conditions, and applicable law.