03 · What You Need to Know
Qualitative anonymisation is not simply deleting identifiers
Context is often part of the qualitative evidence
Qualitative research frequently tries to understand experiences in context. A statement may mean something different depending on who said it, where an event occurred, what relationships surrounded it, and what happened before or afterward.
That creates an unusual anonymisation problem. In a structured dataset, a variable may sometimes be removed with limited effect on the remaining analysis. In an interview narrative, removing one contextual detail can change how the entire account is understood.
UK Data Service guidance notes that text data such as interviews, focus groups, diaries, and fieldnotes contain rich contextual accounts and that this richness increases the possibility of identification through narrative detail. Effective anonymisation therefore requires contextual judgment rather than mechanical deletion.
The identifying detail may also be the explanatory detail
Imagine a participant explaining why they cannot leave their job:
"I am the only speech therapist serving three islands, so if I leave there is nobody for the children."
The occupation and geographic context may help identify the participant. Yet changing the statement to "I have an important job, so if I leave there may be problems" removes much of what makes the account analytically meaningful. Scarcity, geography, professional role, and responsibility are not decorative details. They explain the participant's constraint.
Saunders, Kitzinger, and Kitzinger describe this broader problem in their work on anonymising interview data. Their study required context-sensitive strategies that attempted to maximize anonymity while maintaining the integrity and richness of highly sensitive qualitative accounts.
More anonymisation is not automatically better anonymisation
Removing every contextual characteristic may reduce disclosure risk, but it can produce data that no longer support the analysis for which they were collected.
Anonymisation should therefore be judged partly by whether the resulting material remains useful and truthful. A transcript in which every occupation becomes "[job]," every place becomes "[location]," every relationship becomes "[person]," and every distinctive event becomes "[event]" may technically conceal a great deal while explaining remarkably little.
The problem is particularly acute when the research question concerns place, identity, inequality, organizational roles, relationships, or trajectories over time. Those dimensions cannot always be removed without changing the phenomenon under study.
First identify which details actually create the disclosure risk
Before removing context wholesale, determine what makes the participant identifiable. Sometimes one highly specific detail creates most of the risk while several broader contextual characteristics can safely remain.
For example, an exact hospital name may be unnecessary while the distinction between rural and urban healthcare remains analytically important. An exact job title may be identifying while the broader professional category preserves the relevant interpretation. An exact age may become an age range.
UK Data Service guidance recommends considering contextual and indirect identifiers such as rare occupations, distinctive career trajectories, detailed timelines, highly specific locations, unique personal experiences, and identifiable third parties. It also emphasizes that identification frequently emerges from combinations rather than a single detail.
The aim is therefore not simply to locate every specific detail. It is to determine which details, alone or together, create a realistic route to identification.
Generalization can preserve the analytical category while reducing precision
When exact specificity is not essential, generalization may offer a useful compromise.
Exact detail
"I work at the only neonatal intensive care unit in [small named city]."
Generalized detail
"I work in specialist neonatal care in a regional hospital."
The generalized version removes geographic and institutional precision while retaining the fact that the participant works in specialized hospital care. Whether that is sufficient depends on the analytical question.
Generalization should remain truthful. Replacing a rural hospital with an urban clinic merely because the latter is less identifying is not generalization. It changes the underlying circumstances.
Remove information that does not carry the analytical burden first
A narrative may contain several identifying clues even though only one or two matter to the analysis. Start with the clues that contribute little to interpretation.
Suppose a participant describes a conflict with a supervisor and mentions the supervisor's name, the exact year, the office floor, the participant's department, and the fact that the supervisor controlled promotion decisions. If the analysis concerns hierarchical power, the supervisor's authority may be essential while the name, floor, and exact year are not.
Selective anonymisation can therefore preserve the relationship that explains the finding while removing incidental specificity.
Changing a detail is different from removing or generalizing it
If anonymisation requires modifying a participant's account, researchers need to consider whether the modification alters the evidence.
Replacing a specific municipality with "a rural municipality" may preserve a relevant geographic distinction. Replacing it with a different municipality creates a factual substitution. Similarly, changing a rare occupation into an unrelated occupation may conceal identity while changing why the participant experienced the phenomenon.
The question of whether researchers should alter identifying details inside qualitative quotations therefore requires more than asking whether the change makes identification harder.
Paraphrase can sometimes preserve a finding better than a heavily altered quotation
If a direct quotation cannot be adequately de-identified without substantial modification, paraphrasing may be more transparent.
A researcher can report the analytical substance of an account without presenting reconstructed wording as though it were verbatim speech. This may be preferable when multiple identifying details need to be removed or when the syntax of the original quotation makes substitutions awkward.
Paraphrasing does not automatically eliminate disclosure risk. The underlying event, occupation, or relationship may still identify the participant. It simply gives the researcher more flexibility to summarize the relevant evidence honestly.
Another quotation may support the same finding with less disclosure risk
Not every analytically useful quotation needs to appear in the final publication.
If several participants express the same theme and one quotation is unusually identifying, another excerpt may support the interpretation with less confidentiality risk. This does not mean selecting data merely because they are convenient. The replacement still needs to represent the analytical claim accurately.
The most memorable quotation is sometimes memorable for exactly the same reason it is identifying.
Restricted access can sometimes preserve richer data than public release
Anonymisation is not the only confidentiality safeguard available. The amount of detail that can appropriately remain may depend on who receives the data and under what conditions.
A public dataset or openly accessible publication generally creates a different disclosure environment from a controlled repository, secure research environment, or restricted research team. Contemporary anonymisation guidance from the Information Commissioner's Office treats identifiability as contextual and emphasizes both singling out and linkability when evaluating whether people remain identifiable.
If stripping enough information for unrestricted release would destroy the dataset's research value, controlled access may sometimes provide a more appropriate balance. The feasibility of that approach depends on consent, ethics approval, repository arrangements, applicable law, and institutional requirements.
The audience changes the confidentiality calculation
Kaiser argues that qualitative confidentiality dilemmas can be better understood by considering the audience for the research. Rich detail may be unidentifiable to most readers yet obvious to people who know the participant or setting.
This means anonymisation should not be assessed solely from the researcher's perspective. Ask who will encounter the material and what background information they may possess.
A job title that is safe in a national report may be identifying in an internal organizational report. A village-level location that is acceptable in a large survey may identify the only participant with a particular occupation in a small qualitative sample.
Sometimes the safest choice is not to publish the detail
There are situations in which no satisfactory anonymisation preserves both confidentiality and meaning. The identifying characteristic may be inseparable from the substantive finding.
Researchers should not solve this by pretending the conflict does not exist. Depending on the study, options may include reporting the finding without the specific quotation, combining evidence at a higher level of abstraction, using another example, restricting access, or omitting the material from unrestricted dissemination.
Watch Out
Do not keep removing context until the data become safe but analytically false. If confidentiality cannot be protected without materially changing the interpretation, reconsider the form of disclosure rather than disguising the loss of meaning.