03 · What You Need to Know
Scraping changes the method even when it does not change the source
Web scraping is a collection method, not an ethical category
Web scraping generally refers to automated retrieval and extraction of information from websites or web-accessible systems. Researchers may use simple scripts, browser automation, dedicated crawlers, APIs, or other tools to collect structured or unstructured material.
Whether the resulting research is ethically acceptable depends on what is collected, from where, at what scale, about whom, and for what purpose.
A scraper collecting public government tables for economic analysis raises different concerns from one compiling sensitive person-level profiles from public social media accounts.
The method itself therefore does not automatically make research ethical or unethical.
Public accessibility remains relevant
Researchers should not pretend that there is no difference between scraping public pages and bypassing a private access boundary.
U.S. advisory guidance on internet research distinguishes information legally available to any internet user without special authorization from information available only through a person's permission or an access mechanism under that person's control. The former may be considered public for the relevant regulatory analysis, while the latter may be private.
UKRI similarly states that information intentionally made public in internet spaces may be considered in the public domain, while cautioning that researchers should critically examine the public nature of online information.
The ethical analysis of scraping should preserve this distinction rather than treating all web-accessible material as equivalent.
Technical possibility is not the same as ethical permission
A website may expose information in HTML that a script can retrieve. That establishes a technical fact.
It does not by itself answer whether automated collection complies with applicable terms, licences, access restrictions, privacy and data-protection law, copyright law, institutional ethics requirements, or community expectations.
Technically possible
The researcher can configure software to retrieve the information.
Ethically and legally defensible
The proposed collection method, data scope, and subsequent use satisfy the relevant ethical, institutional, contractual, and legal requirements.
A successful HTTP request is not, alas, a research ethics opinion.
Automation can turn scattered public traces into a person-level dataset
One of the central ethical differences introduced by scraping is aggregation.
Imagine that an individual has hundreds of public posts distributed across several years. Each post reveals relatively little. Automated collection can combine those records into a longitudinal profile of interests, relationships, routines, locations, opinions, and behavior.
HHS advisory guidance on internet research specifically notes that online data permit mining and matching and that partial identifiers can be combined across datasets, allowing individuals to be recognized or revealing surprising new information.
This is why large-scale collection can create ethical problems that are not visible when researchers inspect one public item at a time.
Scraping can collect much more than the research question requires
Automated tools make overcollection remarkably easy. A researcher interested in post text may accidentally or routinely retain usernames, profile photographs, follower counts, timestamps, geolocation, URLs, reactions, network relationships, embedded media, and metadata.
Researchers should define necessary variables before collection rather than treating everything exposed by the page as potentially useful.
The Association of Internet Researchers' guidance encourages researchers to consider whether data collection is necessary and proportionate to the research aim and whether unnecessary information should be deleted.
Data minimization is particularly important in scraping because the marginal technical cost of collecting another field can approach zero while its privacy cost does not.
Scraping public data can still create identifiability
A dataset does not become anonymous because researchers scraped it without collecting real names.
Usernames, profile URLs, images, quotations, locations, network relationships, timestamps, and combinations of attributes may allow identities to be readily determined. HHS guidance emphasizes that identifiability can arise through linkage and association, not merely through the presence of a name.
Researchers should therefore assess the resulting dataset rather than only the source pages.
Scraping can reconstruct sensitive attributes
Automated analysis may infer information that users never explicitly disclosed.
Patterns across posts, follows, locations, vocabulary, interactions, or browsing traces can sometimes support inferences about health, politics, religion, relationships, employment, socioeconomic circumstances, or other sensitive characteristics.
That means the ethical sensitivity of a dataset is not limited to the sensitivity of the fields originally scraped.
The broader distinction between public accessibility and ethically appropriate use becomes particularly important once research generates new information from public traces.
Collection can affect the website or community itself
Scraping is not always passive from the perspective of the source system.
High request volumes can consume bandwidth, trigger rate limits, degrade performance, generate costs, distort analytics, or resemble malicious traffic. Researchers should design collection so that it does not impose unnecessary technical burdens.
This may involve limiting request frequency, collecting only necessary pages, using authorized APIs when appropriate, caching responsibly, and avoiding repeated requests for information already obtained.
The relevant technical safeguards depend on the source and collection architecture.
Access controls should not be treated as puzzles to defeat
Researchers should distinguish public collection from attempts to circumvent authentication, paywalls, CAPTCHAs, rate limits, robots restrictions, or other technical controls.
The legal implications of such actions vary by jurisdiction and system. Ethically, deliberate circumvention also provides evidence that the data holder has attempted to impose a boundary on automated or unrestricted access.
Researchers should not describe information obtained by defeating meaningful access restrictions as ordinary public web data merely because the information eventually appeared on their screen.
Platform rules form another layer of the analysis
UKRI advises researchers using social media to comply with regulations established by data producers where those regulations are consistent with legal and ethical guidance.
Terms of service, API conditions, robots directives, licences, and other platform rules do not themselves constitute a complete research ethics framework. They may nevertheless create contractual, legal, practical, or ethical obligations that researchers need to examine.
Ethics approval does not automatically authorize violation of a platform agreement, just as compliance with platform rules does not automatically make a research design ethically appropriate.
Scraping closed communities raises a different level of concern
A script operating through an authenticated researcher's account may technically be able to collect every post in a private group.
That does not make the material public.
The considerations involved in using restricted online community data continue to apply, and automated extraction may increase the privacy implications by turning a bounded discussion space into a persistent external dataset.
Security becomes more important as scraping increases scale
A scraped dataset may contain thousands or millions of records, including identifiers that researchers do not need for final analysis.
The consequences of unauthorized access can therefore be much larger than in small manual datasets.
Researchers should consider separating identifiers, restricting access, encrypting sensitive files where appropriate, defining retention periods, recording provenance, and deleting unnecessary raw fields once they are no longer required.
Reproducibility does not always require publishing the raw scraped dataset
Open-science norms can create pressure to share data. But redistributing a scraped person-level dataset may expose information in a form far easier to search and analyze than the original website.
Researchers should distinguish transparency about methods from unrestricted redistribution of raw data.
Depending on the study, reproducibility may be supported through code, data dictionaries, derived aggregates, synthetic examples, controlled-access data, or other approaches without publishing a complete identifiable copy of the source.
Watch Out
Do not assume that because every individual record can be found publicly, repackaging millions of those records into a downloadable research dataset creates no additional exposure.