Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Should You Do When Removing Identifying Details Changes the Meaning of Qualitative Data?

Anonymising qualitative data can sometimes remove the very context needed to interpret it. Researchers may need to balance confidentiality with analytical integrity rather than simply deleting every identifying detail.

300
When Anonymisation Changes Meaning Guide 300 of 398
01 · The Question

What if the identifying detail is also part of the explanation?

Removing a participant's name is usually straightforward. Removing everything that could indirectly identify them can be much harder.

Suppose a participant's rural location explains why they could not access specialist healthcare. Their rare occupation explains why workplace pressure affected them differently. Their relationship to another participant explains a power imbalance. Their exact chronology explains how an institutional decision affected them.

If you remove those details, you may protect confidentiality while weakening or even changing the interpretation. Qualitative researchers sometimes face precisely this tension because contextual richness is both analytically valuable and potentially identifying.

02 · The Short Answer

Do not protect identity by silently destroying the analytical meaning

In Brief

When removing identifying details would materially change the meaning of qualitative data, do not automatically strip away the context. Assess whether you can reduce identification risk through selective generalization, omission of less important details, alternative quotation or paraphrase, restricted access, or another disclosure strategy while preserving the information necessary for interpretation.

Sometimes confidentiality and analytical richness genuinely conflict. The appropriate response depends on what was promised to participants, how identifying the information is, why the detail matters analytically, who will receive the data, and whether another form of reporting can preserve the finding without exposing the participant.

03 · What You Need to Know

Qualitative anonymisation is not simply deleting identifiers

Context is often part of the qualitative evidence

Qualitative research frequently tries to understand experiences in context. A statement may mean something different depending on who said it, where an event occurred, what relationships surrounded it, and what happened before or afterward.

That creates an unusual anonymisation problem. In a structured dataset, a variable may sometimes be removed with limited effect on the remaining analysis. In an interview narrative, removing one contextual detail can change how the entire account is understood.

UK Data Service guidance notes that text data such as interviews, focus groups, diaries, and fieldnotes contain rich contextual accounts and that this richness increases the possibility of identification through narrative detail. Effective anonymisation therefore requires contextual judgment rather than mechanical deletion.

The identifying detail may also be the explanatory detail

Imagine a participant explaining why they cannot leave their job:

"I am the only speech therapist serving three islands, so if I leave there is nobody for the children."

The occupation and geographic context may help identify the participant. Yet changing the statement to "I have an important job, so if I leave there may be problems" removes much of what makes the account analytically meaningful. Scarcity, geography, professional role, and responsibility are not decorative details. They explain the participant's constraint.

Saunders, Kitzinger, and Kitzinger describe this broader problem in their work on anonymising interview data. Their study required context-sensitive strategies that attempted to maximize anonymity while maintaining the integrity and richness of highly sensitive qualitative accounts.

More anonymisation is not automatically better anonymisation

Removing every contextual characteristic may reduce disclosure risk, but it can produce data that no longer support the analysis for which they were collected.

Anonymisation should therefore be judged partly by whether the resulting material remains useful and truthful. A transcript in which every occupation becomes "[job]," every place becomes "[location]," every relationship becomes "[person]," and every distinctive event becomes "[event]" may technically conceal a great deal while explaining remarkably little.

The problem is particularly acute when the research question concerns place, identity, inequality, organizational roles, relationships, or trajectories over time. Those dimensions cannot always be removed without changing the phenomenon under study.

First identify which details actually create the disclosure risk

Before removing context wholesale, determine what makes the participant identifiable. Sometimes one highly specific detail creates most of the risk while several broader contextual characteristics can safely remain.

For example, an exact hospital name may be unnecessary while the distinction between rural and urban healthcare remains analytically important. An exact job title may be identifying while the broader professional category preserves the relevant interpretation. An exact age may become an age range.

UK Data Service guidance recommends considering contextual and indirect identifiers such as rare occupations, distinctive career trajectories, detailed timelines, highly specific locations, unique personal experiences, and identifiable third parties. It also emphasizes that identification frequently emerges from combinations rather than a single detail.

The aim is therefore not simply to locate every specific detail. It is to determine which details, alone or together, create a realistic route to identification.

Generalization can preserve the analytical category while reducing precision

When exact specificity is not essential, generalization may offer a useful compromise.

Exact detail "I work at the only neonatal intensive care unit in [small named city]."
Generalized detail "I work in specialist neonatal care in a regional hospital."

The generalized version removes geographic and institutional precision while retaining the fact that the participant works in specialized hospital care. Whether that is sufficient depends on the analytical question.

Generalization should remain truthful. Replacing a rural hospital with an urban clinic merely because the latter is less identifying is not generalization. It changes the underlying circumstances.

Remove information that does not carry the analytical burden first

A narrative may contain several identifying clues even though only one or two matter to the analysis. Start with the clues that contribute little to interpretation.

Suppose a participant describes a conflict with a supervisor and mentions the supervisor's name, the exact year, the office floor, the participant's department, and the fact that the supervisor controlled promotion decisions. If the analysis concerns hierarchical power, the supervisor's authority may be essential while the name, floor, and exact year are not.

Selective anonymisation can therefore preserve the relationship that explains the finding while removing incidental specificity.

Changing a detail is different from removing or generalizing it

If anonymisation requires modifying a participant's account, researchers need to consider whether the modification alters the evidence.

Replacing a specific municipality with "a rural municipality" may preserve a relevant geographic distinction. Replacing it with a different municipality creates a factual substitution. Similarly, changing a rare occupation into an unrelated occupation may conceal identity while changing why the participant experienced the phenomenon.

The question of whether researchers should alter identifying details inside qualitative quotations therefore requires more than asking whether the change makes identification harder.

Paraphrase can sometimes preserve a finding better than a heavily altered quotation

If a direct quotation cannot be adequately de-identified without substantial modification, paraphrasing may be more transparent.

A researcher can report the analytical substance of an account without presenting reconstructed wording as though it were verbatim speech. This may be preferable when multiple identifying details need to be removed or when the syntax of the original quotation makes substitutions awkward.

Paraphrasing does not automatically eliminate disclosure risk. The underlying event, occupation, or relationship may still identify the participant. It simply gives the researcher more flexibility to summarize the relevant evidence honestly.

Another quotation may support the same finding with less disclosure risk

Not every analytically useful quotation needs to appear in the final publication.

If several participants express the same theme and one quotation is unusually identifying, another excerpt may support the interpretation with less confidentiality risk. This does not mean selecting data merely because they are convenient. The replacement still needs to represent the analytical claim accurately.

The most memorable quotation is sometimes memorable for exactly the same reason it is identifying.

Restricted access can sometimes preserve richer data than public release

Anonymisation is not the only confidentiality safeguard available. The amount of detail that can appropriately remain may depend on who receives the data and under what conditions.

A public dataset or openly accessible publication generally creates a different disclosure environment from a controlled repository, secure research environment, or restricted research team. Contemporary anonymisation guidance from the Information Commissioner's Office treats identifiability as contextual and emphasizes both singling out and linkability when evaluating whether people remain identifiable.

If stripping enough information for unrestricted release would destroy the dataset's research value, controlled access may sometimes provide a more appropriate balance. The feasibility of that approach depends on consent, ethics approval, repository arrangements, applicable law, and institutional requirements.

The audience changes the confidentiality calculation

Kaiser argues that qualitative confidentiality dilemmas can be better understood by considering the audience for the research. Rich detail may be unidentifiable to most readers yet obvious to people who know the participant or setting.

This means anonymisation should not be assessed solely from the researcher's perspective. Ask who will encounter the material and what background information they may possess.

A job title that is safe in a national report may be identifying in an internal organizational report. A village-level location that is acceptable in a large survey may identify the only participant with a particular occupation in a small qualitative sample.

Sometimes the safest choice is not to publish the detail

There are situations in which no satisfactory anonymisation preserves both confidentiality and meaning. The identifying characteristic may be inseparable from the substantive finding.

Researchers should not solve this by pretending the conflict does not exist. Depending on the study, options may include reporting the finding without the specific quotation, combining evidence at a higher level of abstraction, using another example, restricting access, or omitting the material from unrestricted dissemination.

Watch Out

Do not keep removing context until the data become safe but analytically false. If confidentiality cannot be protected without materially changing the interpretation, reconsider the form of disclosure rather than disguising the loss of meaning.

04 · A Practical Example

When location is both an identifier and the explanation

Hypothetical Example

A rural participant explains why specialist care was inaccessible

A participant says that receiving treatment required a six-hour boat journey because no specialist service existed on their island. The island is small, and the participant's medical circumstances could make them recognizable locally. Yet geography is central to the research finding about barriers to healthcare access.

Delete everything geographic Reporting only that the participant "had difficulty accessing care" reduces identification risk but removes the mechanism behind that difficulty.
Keep the exact island This preserves maximum context but may make the participant substantially easier to identify.
Generalize strategically The researcher considers whether "a remote island community" preserves the relevant geographic constraint without revealing the exact location.
Review other clues The researcher checks whether diagnosis, age, occupation, family circumstances, or other details would still identify the participant when combined with the generalized location.
Choose the disclosure form If meaningful public reporting still creates unacceptable identification risk, the researcher considers paraphrase, a different example, reduced contextual detail elsewhere, or a more restricted form of access.

The objective is not maximum detail or maximum removal. It is enough context to support the interpretation without retaining unnecessary routes to identification.

05 · What Researchers Often Get Wrong

Common mistakes when anonymisation threatens qualitative meaning

Misconception

The safest dataset is always the best anonymised dataset

A dataset stripped of nearly all context may present less identification risk but also lose much of its analytical value. Effective anonymisation needs to protect participants while preserving enough information for the legitimate research purpose.

Misconception

Every potentially identifying detail should simply be deleted

Some potentially identifying details are also essential explanatory variables or narrative context. Consider whether precision can be reduced, other clues removed, or disclosure controlled before deleting information central to the finding.

Misconception

You can change any fact as long as the participant becomes harder to identify

Changing substantive facts can alter the participant's experience and the interpretation of the evidence. Generalization, omission, paraphrase, or another reporting strategy may be preferable to inventing replacement facts.

Misconception

Anonymisation is finished once names are removed

Qualitative identification often arises from context, specificity, and combinations of details. Names are only the most obvious identifiers.

Misconception

If public release requires destroying the context, the data cannot be shared at all

Not necessarily. Depending on the applicable consent and governance arrangements, controlled or restricted access may sometimes permit richer data to be used while reducing exposure compared with unrestricted release.

06 · What This Means for You

Protect both confidentiality and the integrity of the explanation

When an identifying detail carries analytical meaning, treat the problem as a disclosure decision rather than a simple redaction task.

A simple meaning-preservation framework

If the identifying detail is analytically irrelevant
Remove or mask it according to the study's anonymisation procedures.
If the category matters but the exact value does not
Consider truthful generalization, such as an age range, broader occupation, or less precise geographic description.
If the exact detail carries important analytical meaning
Look for other identifying clues that can be reduced before altering the detail that explains the finding.
If a direct quotation cannot be protected without extensive alteration
Consider paraphrase, another quotation, or another form of reporting rather than presenting heavily reconstructed language as verbatim evidence.
If meaningful anonymisation is incompatible with unrestricted disclosure
Consider whether controlled access or another approved disclosure arrangement can preserve greater analytical utility while managing identification risk.

Document these decisions. An anonymisation log or protocol can record what was changed, the reason for the change, and the principle applied to similar material. That helps prevent confidentiality decisions from becoming a collection of unexplained editorial improvisations.

07 · A Quick Checklist

Before removing an identifying qualitative detail

Check whether the proposed anonymisation:
Targets information that actually contributes to a realistic identification risk.
Removes incidental identifying details before altering information central to the analysis.
Can use a broader truthful category instead of deleting the relevant context completely.
Preserves the mechanism, relationship, chronology, or circumstance necessary to interpret the finding.
Avoids introducing fictional details that materially change the participant's account.
Accounts for combinations of remaining details rather than evaluating each identifier separately.
Considers who will receive the material and what contextual information those recipients may possess.
Considers another quotation, paraphrase, omission, or controlled access when public anonymisation would destroy analytical meaning.
08 · Frequently Asked Questions

Questions about balancing anonymity and qualitative meaning

Should I remove every detail that could potentially identify a participant?

Not mechanically. Assess realistic identification risk, the information's analytical importance, the likely audience, and the safeguards available. Some details may be generalized or managed through other controls rather than completely removed.

Can I generalize a participant's location?

Often, if the broader description remains truthful and retains the geographic distinction needed for analysis. Whether "a remote island community," "a rural province," or another level of generalization is appropriate depends on the study and disclosure risk.

What if occupation is both identifying and important to the finding?

Consider whether a broader occupational category preserves the relevant mechanism, whether other identifying details can be reduced instead, or whether another quotation or disclosure arrangement can communicate the finding without exposing the participant.

Can I paraphrase instead of anonymising a direct quote heavily?

Yes, when appropriate. Paraphrase can preserve the substantive finding without implying that extensively modified wording is verbatim. The paraphrase itself should still be checked for identifying contextual information.

Can restricted access solve the problem?

It can sometimes reduce disclosure risk enough to preserve richer data that would be unsuitable for unrestricted release. Whether controlled access is available and appropriate depends on consent, ethics approval, repository arrangements, institutional requirements, and applicable law.

What if no anonymisation preserves both confidentiality and meaning?

You may need to change the form of disclosure rather than continue modifying the data. Options can include omitting the quotation, reporting the finding at a higher level of abstraction, using another example, or limiting access under an appropriate approved arrangement.

09 · The Bottom Line

Anonymisation should protect participants without hollowing out the evidence

The Bottom Line

When removing identifying details changes the meaning of qualitative data, do not automatically sacrifice the context that makes the finding intelligible. Reduce unnecessary specificity first, preserve analytically essential information where it can be protected appropriately, and change the form or conditions of disclosure when meaningful anonymisation is not possible.

Qualitative confidentiality sometimes requires compromise rather than perfect concealment with perfect data fidelity. The defensible choice is the one that takes both obligations seriously and does not quietly solve a privacy problem by creating an interpretation problem.

10 · Sources and Further Reading

Research and guidance on anonymising qualitative data

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes