Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Does Web Scraping Create Ethical Responsibilities Even When the Data Are Public and Scraping Is Technically Possible?

The ability to scrape publicly accessible information does not settle whether doing so is ethically appropriate. Automated collection can change scale, persistence, identifiability, system burden, and the consequences of combining public data.

361
Web Scraping Research Ethics Guide 361 of 398
01 · The Question

If the data are public and a computer can collect them, what ethical problem remains?

A researcher can manually open a webpage and read its contents. A script can perform the same basic action thousands or millions of times.

It is tempting to view web scraping as nothing more than efficient reading.

Sometimes that comparison is reasonable. But automation can change the scale, speed, persistence, granularity, and consequences of collection. A person browsing profiles individually cannot realistically assemble the same dataset as a system that captures every post, timestamp, username, relationship, location, and profile field across millions of accounts.

The ethical question is therefore not only whether each source page is public. It is also what automated collection makes possible.

02 · The Short Answer

Yes, automation can create responsibilities beyond public accessibility

In Brief

Yes. Web scraping can create ethical responsibilities even when source information is publicly accessible and automated collection is technically possible, because scraping can alter the scale, persistence, identifiability, aggregation, system burden, and downstream uses of the data.

Researchers should assess whether the data and volume are necessary, what people reasonably expect, whether records can be linked or used to infer sensitive information, how collection affects the source system or community, and what platform, contractual, ethics, privacy, copyright, and legal requirements apply.

03 · What You Need to Know

Scraping changes the method even when it does not change the source

Web scraping is a collection method, not an ethical category

Web scraping generally refers to automated retrieval and extraction of information from websites or web-accessible systems. Researchers may use simple scripts, browser automation, dedicated crawlers, APIs, or other tools to collect structured or unstructured material.

Whether the resulting research is ethically acceptable depends on what is collected, from where, at what scale, about whom, and for what purpose.

A scraper collecting public government tables for economic analysis raises different concerns from one compiling sensitive person-level profiles from public social media accounts.

The method itself therefore does not automatically make research ethical or unethical.

Public accessibility remains relevant

Researchers should not pretend that there is no difference between scraping public pages and bypassing a private access boundary.

U.S. advisory guidance on internet research distinguishes information legally available to any internet user without special authorization from information available only through a person's permission or an access mechanism under that person's control. The former may be considered public for the relevant regulatory analysis, while the latter may be private.

UKRI similarly states that information intentionally made public in internet spaces may be considered in the public domain, while cautioning that researchers should critically examine the public nature of online information.

The ethical analysis of scraping should preserve this distinction rather than treating all web-accessible material as equivalent.

Technical possibility is not the same as ethical permission

A website may expose information in HTML that a script can retrieve. That establishes a technical fact.

It does not by itself answer whether automated collection complies with applicable terms, licences, access restrictions, privacy and data-protection law, copyright law, institutional ethics requirements, or community expectations.

Technically possible The researcher can configure software to retrieve the information.
Ethically and legally defensible The proposed collection method, data scope, and subsequent use satisfy the relevant ethical, institutional, contractual, and legal requirements.

A successful HTTP request is not, alas, a research ethics opinion.

Automation can turn scattered public traces into a person-level dataset

One of the central ethical differences introduced by scraping is aggregation.

Imagine that an individual has hundreds of public posts distributed across several years. Each post reveals relatively little. Automated collection can combine those records into a longitudinal profile of interests, relationships, routines, locations, opinions, and behavior.

HHS advisory guidance on internet research specifically notes that online data permit mining and matching and that partial identifiers can be combined across datasets, allowing individuals to be recognized or revealing surprising new information.

This is why large-scale collection can create ethical problems that are not visible when researchers inspect one public item at a time.

Scraping can collect much more than the research question requires

Automated tools make overcollection remarkably easy. A researcher interested in post text may accidentally or routinely retain usernames, profile photographs, follower counts, timestamps, geolocation, URLs, reactions, network relationships, embedded media, and metadata.

Researchers should define necessary variables before collection rather than treating everything exposed by the page as potentially useful.

The Association of Internet Researchers' guidance encourages researchers to consider whether data collection is necessary and proportionate to the research aim and whether unnecessary information should be deleted.

Data minimization is particularly important in scraping because the marginal technical cost of collecting another field can approach zero while its privacy cost does not.

Scraping public data can still create identifiability

A dataset does not become anonymous because researchers scraped it without collecting real names.

Usernames, profile URLs, images, quotations, locations, network relationships, timestamps, and combinations of attributes may allow identities to be readily determined. HHS guidance emphasizes that identifiability can arise through linkage and association, not merely through the presence of a name.

Researchers should therefore assess the resulting dataset rather than only the source pages.

Scraping can reconstruct sensitive attributes

Automated analysis may infer information that users never explicitly disclosed.

Patterns across posts, follows, locations, vocabulary, interactions, or browsing traces can sometimes support inferences about health, politics, religion, relationships, employment, socioeconomic circumstances, or other sensitive characteristics.

That means the ethical sensitivity of a dataset is not limited to the sensitivity of the fields originally scraped.

The broader distinction between public accessibility and ethically appropriate use becomes particularly important once research generates new information from public traces.

Collection can affect the website or community itself

Scraping is not always passive from the perspective of the source system.

High request volumes can consume bandwidth, trigger rate limits, degrade performance, generate costs, distort analytics, or resemble malicious traffic. Researchers should design collection so that it does not impose unnecessary technical burdens.

This may involve limiting request frequency, collecting only necessary pages, using authorized APIs when appropriate, caching responsibly, and avoiding repeated requests for information already obtained.

The relevant technical safeguards depend on the source and collection architecture.

Access controls should not be treated as puzzles to defeat

Researchers should distinguish public collection from attempts to circumvent authentication, paywalls, CAPTCHAs, rate limits, robots restrictions, or other technical controls.

The legal implications of such actions vary by jurisdiction and system. Ethically, deliberate circumvention also provides evidence that the data holder has attempted to impose a boundary on automated or unrestricted access.

Researchers should not describe information obtained by defeating meaningful access restrictions as ordinary public web data merely because the information eventually appeared on their screen.

Platform rules form another layer of the analysis

UKRI advises researchers using social media to comply with regulations established by data producers where those regulations are consistent with legal and ethical guidance.

Terms of service, API conditions, robots directives, licences, and other platform rules do not themselves constitute a complete research ethics framework. They may nevertheless create contractual, legal, practical, or ethical obligations that researchers need to examine.

Ethics approval does not automatically authorize violation of a platform agreement, just as compliance with platform rules does not automatically make a research design ethically appropriate.

Scraping closed communities raises a different level of concern

A script operating through an authenticated researcher's account may technically be able to collect every post in a private group.

That does not make the material public.

The considerations involved in using restricted online community data continue to apply, and automated extraction may increase the privacy implications by turning a bounded discussion space into a persistent external dataset.

Security becomes more important as scraping increases scale

A scraped dataset may contain thousands or millions of records, including identifiers that researchers do not need for final analysis.

The consequences of unauthorized access can therefore be much larger than in small manual datasets.

Researchers should consider separating identifiers, restricting access, encrypting sensitive files where appropriate, defining retention periods, recording provenance, and deleting unnecessary raw fields once they are no longer required.

Reproducibility does not always require publishing the raw scraped dataset

Open-science norms can create pressure to share data. But redistributing a scraped person-level dataset may expose information in a form far easier to search and analyze than the original website.

Researchers should distinguish transparency about methods from unrestricted redistribution of raw data.

Depending on the study, reproducibility may be supported through code, data dictionaries, derived aggregates, synthetic examples, controlled-access data, or other approaches without publishing a complete identifiable copy of the source.

Watch Out

Do not assume that because every individual record can be found publicly, repackaging millions of those records into a downloadable research dataset creates no additional exposure.

04 · A Practical Example

The ethical difference appears when the scraper connects the dots

Hypothetical Example

Scraping public posts about commuting

A research team wants to study how commuters discuss transport disruptions. Relevant posts are publicly visible without login.

Initial plan The team proposes scraping every matching post over two years together with usernames, exact timestamps, profile biographies, profile locations, follower networks, photographs, and all linked posts from each account.
Research need The actual research question requires post text, broad time periods, and disruption type. Most profile and network variables do not contribute to the analysis.
Risk The original plan would permit longitudinal reconstruction of individual commuters' routines and locations despite the study having no person-level research question.
Redesign The team narrows collection to relevant posts and necessary metadata, limits request rates, removes unnecessary identifiers at the earliest appropriate stage, and defines retention and access controls.
Governance The researchers verify institutional ethics requirements, platform conditions, applicable law, and whether the proposed automated access method is permitted before collection begins.

The posts did not become less public. The revised protocol simply stopped collecting information the study never needed.

05 · What Researchers Often Get Wrong

Public pages and ethical scraping are not synonymous

Misconception

If I can scrape it, I am allowed to scrape it

Technical capability establishes only that the software can retrieve the information. Researchers must separately consider ethics, access restrictions, platform conditions, contractual obligations, copyright, privacy, data protection, and other applicable law.

Misconception

Scraping is just automated reading

At small scales it may resemble manual retrieval, but automation can create qualitatively different capabilities through aggregation, longitudinal tracking, linkage, profiling, and collection of entire populations.

Misconception

If every field is public, collecting every field is harmless

Data minimization still matters. Combining individually public attributes can increase identifiability and reveal information that no single field provides.

Misconception

A dataset without real names is anonymous

Usernames, URLs, quotations, networks, locations, timestamps, and combinations of attributes may still permit identification or linkage to identifiable profiles.

Misconception

Open science means raw scraped data should always be published

Research transparency does not require researchers to create avoidable privacy or legal risks. Sharing strategies should reflect identifiability, licences, platform conditions, consent, and the sensitivity of the resulting dataset.

06 · What This Means for You

Design the scraper around the research question, not around everything the website exposes

Before writing collection code, define what variables and volume the study genuinely requires. That decision is part of research ethics, not merely software engineering.

A simple decision framework

If the required information is unrestricted public data
Public status may support collection, but still assess scale, identifiability, platform conditions, proportionality, and downstream use.
If collection requires bypassing a meaningful access or technical restriction
Do not assume the material remains ordinary public data. Examine the ethical and legal significance of the restriction before proceeding.
If the scraper captures fields unrelated to the research question
Exclude them at collection where practical or remove them as early as scientifically appropriate.
If collection enables person-level linkage or sensitive inference
Reassess privacy, identifiability, security, ethics-review status, and whether that level of analysis is actually necessary.
If raw-data sharing would substantially increase exposure
Use a reproducibility strategy proportionate to the data rather than assuming unrestricted redistribution is required.

Efficient collection is a technical achievement. Collecting only what the study can justify is the research achievement.

07 · A Quick Checklist

Before scraping public web data, check the method as well as the source

Before running the scraper, check:
Verify whether the source is genuinely public or requires authentication, permission, payment, membership, or another meaningful access condition.
Define the minimum records, fields, dates, pages, and metadata necessary for the research question.
Assess whether combining records can identify individuals or infer sensitive characteristics not obvious from individual source items.
Design request frequency and collection behavior to avoid unnecessary technical burden on the source system.
Review relevant platform terms, API conditions, licences, robots directives, and technical access restrictions.
Avoid collecting usernames, profile data, images, networks, or other identifiers merely because they are technically available.
Plan secure storage, access, de-identification, retention, and deletion appropriate to the scale and sensitivity of the resulting dataset.
Determine what data can ethically and legally be shared for reproducibility rather than automatically publishing the raw scrape.
Verify applicable institutional ethics, privacy, data-protection, copyright, contractual, platform, and other legal requirements before automated collection begins.
08 · Frequently Asked Questions

Questions about web scraping for research

Is it ethical to scrape publicly accessible websites for research?

Potentially. Public accessibility is relevant, but researchers should also consider proportionality, scale, identifiability, platform conditions, technical burden, data security, ethics requirements, and applicable law.

Does robots.txt determine whether research scraping is ethical?

No single technical file determines the complete ethics analysis. Robots directives can provide relevant information about automated access expectations, but researchers must also consider platform conditions, licences, law, research ethics, privacy, and the actual effects of collection.

Can I scrape data behind a login if I have an account?

Account access does not automatically make restricted information public research data. Researchers should examine why access is restricted, the terms governing the account, user expectations, consent, ethics review, and applicable law.

Do I need ethics review if all scraped data are public?

Requirements depend on the governing framework and research design. Some research using publicly available information may be exempt or outside formal human-subjects review, but researchers should obtain the appropriate institutional determination where status is uncertain and still address broader ethical responsibilities.

Should researchers remove usernames immediately after scraping?

When usernames are unnecessary, early removal can reduce risk. Some studies legitimately require account-level linkage, in which case researchers should justify retaining identifiers and protect them appropriately rather than removing them reflexively.

Can researchers publish a scraped dataset if all records were originally public?

Not automatically. Redistribution can substantially increase accessibility, searchability, linkage, and persistence. Researchers should check identifiability, licences, platform conditions, consent where relevant, ethics commitments, copyright, and applicable law before sharing raw data.

09 · The Bottom Line

The ability to automate collection does not automate the ethical judgment

The Bottom Line

Web scraping can create ethical responsibilities even when every source item is publicly accessible, because automated collection can transform scattered information into large, persistent, linkable, and potentially identifying datasets.

Ask not only whether the website exposes the data, but why you need each field, how much you need, what the combined dataset reveals, what burden collection creates, and what rules govern access and reuse. The scraper should implement an ethically justified research design, not define one by whatever it happens to be capable of downloading.

10 · Sources and Further Reading

Authoritative guidance relevant to automated internet research

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes