01 · The Question
If anyone can see a post, is it automatically public research data?
A social media user posts something from an unrestricted account. No login, membership, or permission is needed to see it. A researcher can find the post through the platform itself or perhaps through a search engine.
Does that settle the ethics question?
Not quite. Public accessibility is important, and some research frameworks explicitly recognize genuinely public online information differently from private information. But research can do things that an ordinary viewer does not: collect thousands of posts, preserve deleted material, connect profiles, infer characteristics, reproduce quotations, and move someone's words into an entirely different audience and context.
The useful question is therefore not only whether the post is public. It is what “public” establishes, and what it does not.
03 · What You Need to Know
The word “public” answers only part of the research ethics question
Some ethics frameworks do recognize genuinely public online information
Researchers should not swing too far in the opposite direction and assume that everything written by a person online must be treated as confidential participant data.
For example, UK Research and Innovation's Economic and Social Research Council guidance states that information in internet spaces intentionally made public would be considered in the public domain. The same guidance immediately cautions that the public nature of internet and social media communications should always be critically examined.
This distinction is useful. Public accessibility can materially affect whether individual consent or ethics review is required. It simply does not resolve every downstream decision about collection and use.
Platform visibility is evidence about publicity, not the entire definition
A post from an unrestricted account differs from content visible only to approved followers. A message on an open broadcast channel differs from one inside an invitation-only support community.
Those technical distinctions matter because they indicate who can ordinarily access the content. But platforms also have norms, audiences, architectures, and user practices that influence what posting publicly means in context.
Recent methodological scholarship, for example, has distinguished between social media content functioning more like intentionally public creative production and content functioning more like sensitive conversation conducted in a technically open environment. Researchers have argued that the platform's genre, norms, and affordances can therefore affect the appropriate ethical treatment of apparently public material.
A post can be public while the person behind it remains ethically relevant
Calling something public data can subtly shift attention from the person who produced it to the data object itself. That shift is sometimes methodologically appropriate, but researchers should not allow the terminology to erase potential effects on the author.
A public post can reveal health conditions, financial difficulties, political activity, religious affiliation, sexuality, experiences of violence, workplace disputes, illegal behavior, or information about other people. Research reuse may expose those disclosures to an audience dramatically different from the one that originally encountered them.
The Association of Internet Researchers recommends examining whether data can identify people directly or through inference, whether sensitive characteristics can be reconstructed, whether collected information is necessary and proportionate to the research aim, and whether unnecessary data should be deleted.
The broader question of whether internet users should be treated as participants or data sources therefore cannot always be answered simply by looking at a platform's privacy icon.
Public posting and consent to research are different concepts
A user who intentionally publishes something publicly has made the content available to an unrestricted audience. That fact may be highly relevant to the ethical and regulatory status of research using it.
It does not follow that the person has completed an informed-consent process for every possible research project involving the material.
Making a post public
The user makes content available to an audience defined by the platform's access configuration and the circumstances of publication.
Consenting to a research project
A person receives relevant information about a particular study and voluntarily agrees to participate under the applicable consent process.
Research may sometimes legitimately use public information without obtaining such consent. The justification, however, is not that public posting secretly functions as a universal research consent form.
Reasonable expectations of privacy can survive some forms of public accessibility
Online privacy is not always binary. Research has long noted that some users of open social media do not experience their communication as equivalent to publishing material for unlimited institutional analysis, and automated searching and aggregation can turn scattered public traces into structured dossiers.
UKRI consequently advises researchers to critically examine the public nature of online communication and specifically warns that some social media users, including children, may not understand the implications of what they are doing.
This does not mean researchers should assume that every public post is secretly private. It means reasonable expectations should be assessed rather than inferred mechanically from one technical setting.
Sensitivity can matter even when publicity is unambiguous
Suppose two users deliberately publish unrestricted posts. One announces the release of a new book. The other describes a recent suicide attempt under a pseudonym.
Both posts may be equally accessible at the platform level. Their ethical implications for research reuse are not necessarily equivalent.
Sensitivity affects the consequences of identification, quotation, aggregation, and unexpected exposure. Research on public social media data has consequently argued against treating “public” and “private” as a complete binary for ethical decision-making.
Removing the username may not anonymize a public post
Social media text is searchable. If a researcher reproduces a distinctive sentence verbatim, a reader may be able to enter the quotation into a search engine or platform search and find the original post.
The author can then become identifiable even though the paper contains no username.
This is why online quotations can identify their authors and why decisions about quoting social media posts deserve separate attention from the initial decision to collect the data.
Paraphrasing may protect identity, but it can also change the data
Researchers sometimes reduce searchability by paraphrasing online material rather than reproducing exact wording. That can be useful when semantic content matters more than precise language.
It may be inappropriate when wording, rhetoric, discourse, spelling, syntax, or other linguistic features are themselves objects of analysis. The decision about whether to paraphrase online posts to reduce searchability therefore involves both ethical and methodological considerations.
Scale can transform what “using a public post” means
There is a meaningful difference between analyzing a handful of public statements and automatically collecting millions of posts together with usernames, timestamps, locations, networks, profile attributes, and interaction histories.
Each item may remain public. The resulting dataset may nevertheless enable inferences and forms of surveillance that no individual post could support.
AoIR's ethical guidance specifically asks researchers whether data collection is necessary and proportionate, whether people can be identified indirectly through inference, and whether sensitive characteristics can be reconstructed from collected data.
This is why large-scale collection can create ethical problems even when individual records appear innocuous.
Public today does not necessarily mean public forever
Users delete posts, change usernames, restrict accounts, remove photographs, and close profiles. Researchers may still possess material collected while it was public.
The ethics of retaining and publishing content that was deleted after collection therefore cannot be reduced to its status at the instant it entered the dataset.
Researchers should consider the purpose of retention, potential harm, commitments made in the protocol, applicable law, and whether continued use remains proportionate.
04 · A Practical Example
Two public posts can require different research decisions
Hypothetical Example
Studying public reactions to workplace conditions
A researcher searches a public social media platform for posts about workplace experiences. The search returns posts from accounts that anyone can view without following the authors.
Post A
A verified professional account publishes a formal public statement about working conditions and explicitly addresses journalists, policymakers, and the general public.
Post B
A pseudonymous individual describes harassment by a supervisor and includes distinctive details about the employer, colleagues, location, and incident.
Accessibility
Both posts are publicly visible under the platform's technical settings.
Ethical difference
The second post presents substantially greater risks of identification and harm if reproduced or linked to its author. Its conversational and sensitive context may also warrant stronger protection.
Research decision
The researcher evaluates collection and publication separately, considering whether exact quotation or identifying metadata are necessary for the analysis rather than treating the shared “public” setting as sufficient justification for identical handling.
Publicity matters. Context determines what researchers should responsibly do with it.
07 · A Quick Checklist
Before treating public social media posts as research data, check the context
Before collecting public posts, check:
Verify that the content is genuinely available to an unrestricted audience rather than accessible only through membership, approval, credentials, or another access condition.
Consider whether the platform, account, community, and type of post indicate an intention to communicate broadly or primarily within a particular social context.
Identify sensitive personal information and the consequences if the post is linked back to its author.
Test whether usernames, exact quotations, images, timestamps, locations, or contextual details make the author identifiable.
Collect only profile information, metadata, posts, and contextual details necessary for the research question.
Separate the justification for collecting a public post from the decision to quote, reproduce, archive, or publish it.
Assess whether automated collection, linkage, or aggregation creates risks beyond those presented by individual posts.
Verify applicable ethics-review requirements, platform conditions, privacy and data-protection law, and other institutional requirements.