03 · What You Need to Know
Online research blurs the boundary between public information and human participation
Start by asking what kind of online environment you are studying
“The internet” is not one research setting. A government website, unrestricted discussion board, public social media profile, pseudonymous support forum, password-protected community, private messaging channel, and invitation-only group create very different relationships between users and audiences.
Researchers should therefore avoid beginning with the blanket question, “Is internet data public?” The more useful starting point is: what is this particular environment, who can ordinarily access it, what restrictions exist, and what audience could contributors reasonably have had in mind?
Public accessibility can matter without ending the analysis
Canada's Panel on Research Ethics guidance for research using social media recognizes that existing social media information may be either private or in the public domain and that the boundary is not always clear. It notes that genuinely unrestricted content intentionally made available publicly may be considered accessible to others, including for research purposes. At the same time, the guidance emphasizes a spectrum of access and directs researchers to consider user intent, privacy settings, and other contextual factors.
This is why the question of whether a public social media post automatically counts as public research data requires more than checking whether the post appears without a login.
Technical accessibility and ethical publicity are different questions
A researcher may technically be able to retrieve information while still having reasons to treat it cautiously.
Technically accessible
The researcher can retrieve the information using a browser, search engine, API, scraper, account, archive, or other means.
Ethically appropriate to use
The proposed collection, analysis, linkage, quotation, and dissemination are defensible given the context, expectations, risks, and applicable research requirements.
The later distinction between publicly accessible and ethically public information becomes especially important in online environments because digital information can travel far beyond the audience for which it was originally produced.
User intent can be ambiguous
Privacy settings provide useful evidence about intended audience, but they do not always tell the entire story.
Some users deliberately publish content to reach the widest possible audience. Others may accept default settings without understanding them. People may post pseudonymously in a technically open forum while experiencing the forum as a bounded community. A message may be publicly viewable but buried deep enough that contributors reasonably anticipate an audience composed mostly of other community members.
The Canadian social media guidance specifically recognizes that users may or may not understand the implications of their privacy settings and that platforms provide different degrees of audience control.
Researchers should therefore avoid treating a platform's interface setting as a perfect proxy for a person's understanding of research reuse.
Researchers need to consider context, not just content
An individual may voluntarily disclose sensitive information online without expecting it to become part of a research dataset.
Consider a pseudonymous forum where people discuss infertility, addiction, workplace harassment, immigration problems, financial distress, or experiences of violence. Even when individual posts can technically be viewed by outsiders, systematic collection may create risks not obvious from reading a single post.
Researchers should ask what could happen if the content were linked back to its author, community, location, or other contextual information.
Pseudonyms do not necessarily make people anonymous
A username is not the only route to identification. Searchable text, distinctive phrases, dates, biographical details, images, linked profiles, locations, and combinations of seemingly harmless facts may lead back to the original author.
This creates a particular problem for qualitative internet research. A researcher may remove the username but reproduce a sentence verbatim. A reader can paste that sentence into a search engine and potentially find the original post and profile.
Whether researchers need permission to quote social media posts, and how quoting online content can make an author identifiable, therefore require separate consideration.
Paraphrasing can sometimes reduce searchability
When the precise wording is not analytically necessary, researchers may consider paraphrasing material so that a search engine cannot easily connect the published research back to the source post.
That strategy involves trade-offs. Paraphrasing may reduce identification risk, but it can also alter tone, meaning, discourse features, or linguistic phenomena that matter to the analysis. Researchers therefore need to decide whether paraphrasing online posts to reduce searchability is compatible with the research question.
Closed and private spaces require different reasoning
Information accessible only after joining a group, receiving approval, using credentials, paying for membership, or being accepted by a moderator should not automatically be treated like unrestricted public content.
Researchers should examine data from closed or private online communities with particular attention to access conditions, community expectations, researcher disclosure, consent, and gatekeeper authority.
Simply obtaining technical access does not necessarily answer the ethical question. Likewise, joining a private online group does not automatically grant permission to convert members' interactions into research data.
Internet users can occupy more than one ethical role
Researchers sometimes describe online material as documentary data rather than information obtained from human participants. In other projects, researchers interact with users, follow individuals over time, solicit responses, intervene in communities, or construct detailed profiles from their digital traces.
The ethical relationship can therefore vary. The question of whether internet users are research participants or data sources should be examined in relation to the specific method rather than answered by platform type alone.
Scale can change the ethical character of collection
Reading ten publicly visible posts manually is not necessarily equivalent to scraping ten million posts and combining them with location, profile, network, or behavioral information.
Aggregation can reveal patterns or attributes that no individual data item discloses. Large datasets can also increase the number of people potentially affected by re-identification, misuse, security failures, or unexpected secondary analyses.
The Association of Internet Researchers' Internet Research: Ethical Guidelines 3.0 encourages researchers to ask whether collected data are necessary and proportionate to the research aim, whether individuals can be identified directly or indirectly, whether sensitive characteristics can be reconstructed, and whether unnecessary information is deleted.
Questions about large-scale data collection becoming ethically problematic therefore extend beyond the sensitivity of any single post.
Web scraping adds another layer of responsibility
Automated collection can dramatically expand the volume, speed, persistence, and granularity of online research data. It can also implicate platform rules, technical restrictions, data-protection law, security, and the burden placed on websites or communities.
The fact that a researcher can technically retrieve information at scale does not settle whether web scraping of public data is ethically appropriate.
Deleted content creates a temporal privacy problem
Online information can change after collection. A user may delete a post, close an account, change privacy settings, or remove material they regret sharing.
Researchers may still possess a copy collected while the material was visible. That creates a difficult question about whether technically legitimate collection at one moment justifies continued retention, quotation, or publication after the author's apparent decision to withdraw the material from public view.
The treatment of deleted online content that was public when collected therefore requires its own contextual judgment.
Platform terms, research ethics, and law are separate layers
Researchers may need to consider several overlapping forms of permission and restriction: research ethics requirements, privacy and data-protection law, intellectual property, contractual platform terms, API conditions, and community rules.
Compliance with one layer does not automatically satisfy all the others. A platform may technically provide data access while an ethics committee still identifies privacy concerns. Conversely, an ethics approval does not grant researchers immunity from applicable law or contractual restrictions.
Watch Out
Do not write “the data are public, therefore no ethics issues apply” as the entire justification for internet research. Explain how the information became accessible, what users could reasonably expect, what you will collect, how people could be identified, and what your research will do with the material.
04 · A Practical Example
A public discussion forum can still create identification risks
Hypothetical Example
Studying experiences discussed in an open online forum
A researcher wants to analyze how people describe difficult workplace experiences. A discussion forum can be read without registering, and hundreds of relevant posts appear in search results. Most contributors use pseudonyms.
Access
The posts are technically available to anyone using an ordinary web browser, which supports treating accessibility as an important part of the ethical assessment.
Context
The researcher examines whether contributors appear to use the forum as a broad public publishing space or as a community centered on peer discussion despite its open technical configuration.
Sensitivity
Some posts contain detailed accounts of employers, health problems, relationships, locations, and conflicts that could create harm if linked back to contributors.
Identification
Although usernames can be removed, verbatim quotations may be searchable and lead readers directly to the original posts.
Research design
The researcher minimizes unnecessary profile data, determines when quotation is analytically necessary, considers paraphrasing where appropriate, and establishes safeguards for storing the collected material.
Ethics determination
The researcher follows the institution's applicable review process rather than assuming that browser access alone settles whether consent or additional protections are required.
The ethical question is therefore larger than “Could I read this post?” Research transforms reading into systematic collection, analysis, retention, and dissemination.
07 · A Quick Checklist
Before using online content without consent, examine the context
Before collecting online data, check:
Determine whether the content is genuinely unrestricted or requires an account, membership, approval, credentials, or other access condition.
Consider what audience users appear to have intended and whether community norms create expectations beyond the platform's technical settings.
Assess whether the content contains sensitive personal information or involves people who could experience disproportionate harm from identification.
Test whether usernames, quotations, images, timestamps, locations, or combinations of details could lead readers back to individual users.
Collect only the posts, profile fields, metadata, and contextual information genuinely needed for the research question.
Decide how quotations will be reported and whether exact wording creates unnecessary searchability.
Assess whether automated or large-scale collection creates risks not apparent from individual posts.
Check applicable platform conditions, community rules, privacy and data-protection law, copyright considerations, and institutional ethics requirements.
Document why using the material without individual consent is ethically appropriate for this specific dataset and research design.