AI Thematic Analysis Flattens Vulnerable Voices: 2026 Study
Qualitati Research Team · 2026-08-26 · 6 min read
AI thematic analysis can reproduce a study's descriptive themes while quietly softening its most politically loaded ones. In a 2026 comparative study, an LLM re-analysing interviews with disabled students matched human themes above chance but substituted participants' own words — replacing "ableism" with "discrimination" — and generalised first-person disclosures into claims about users.
What did the study test?
The study compared human-generated and LLM-generated thematic analysis of the same interview corpus. Markelius, Dogan, Bailey and Gunes (2026), in a paper accepted to the 35th IEEE RO-MAN conference, analysed semi-structured interviews with 31 university students (mean age 25.2, SD = 6.7) who had disclosed a disability: autism (n = 5), specific learning differences (n = 7), or a mental health condition (n = 7). The interviews came from a 2×2 within-subjects human-robot interaction study on robot versus voice-agent support.
The authors ran the LLM through Braun and Clarke's six-phase reflexive thematic analysis, with the model given a system prompt casting it as a thematic analyst and supplied with study context. Familiarisation and coding were done in four batches to manage token load; theme generation, review and naming were done over the pooled codebook. The reported model was Claude Sonnet 4.6 with extended thinking enabled.
How closely did the AI themes match the human themes?
Above chance, but not close. According to Markelius et al. (2026), adjusted mutual information (AMI) between the human and LLM clusterings was 0.231 at the theme level and 0.225 at the subtheme level. Permutation testing confirmed both were well above chance — but an AMI in the low 0.2s describes a partial, not a substitutable, overlap.
Semantic agreement, measured as cosine similarity over all-mpnet-base-v2 sentence embeddings, showed a clear gradient: similarity was highest at the code level and decreased as the analysis moved up to subthemes and themes. In other words, the AI agreed most where the work was closest to the transcript, and least where interpretation did the heavy lifting.
Where agreement was highest and lowest
| Analytic level or theme | Agreement with human analysis | What it suggests |
| Code level | Highest semantic similarity | Close-to-transcript labelling transfers well |
| Subtheme level | Lower | Grouping decisions start to diverge |
| Theme level | Lowest | Abstraction is where the models and humans part ways |
| "Sensitivity to disability needs" | Highest of the themes | Descriptive, surface-level content is reproducible |
| "Power dynamics" / "effort requirements" | Lowest of the themes | Socially situated concepts resist automation |
The human theme "power dynamics" was distributed roughly evenly across three separate LLM categories — the model had no single home for it.
What went wrong, specifically?
The failure mode was not hallucination. The authors report no fabricated content; the divergences were systematic and interpretive. Four patterns stand out:
- Paraphrase instead of interpretation. The model produced extractive or pattern-generalised codes that, in the authors' framing, do not sit with the utterance in depth.
- Semantic sanitisation. A participant's explicit term — "ableism" — was rendered as the blander, structurally neutral "discrimination". The referent survives; the political claim does not.
- Loss of specificity. A first-person disclosure of lived anxiety became a generalised statement about anxious users, moving the account from a person to a user segment.
- Weak abstraction. The model handled descriptive themes competently and struggled with abstract, socially situated ones.
The authors read this as an epistemic problem rather than a bug: LLMs may default to hegemonic framings, which is a benign tendency in a product-feedback corpus and a consequential one when the corpus is about identity, power and structural exclusion.
What this means for researchers
The practical reading is a division of labour, not a verdict. The paper's own recommendation is that LLMs are appropriate for early-stage work — coding and clustering — with human-in-the-loop review reserved for higher-level interpretation, and that hybrid workflows should have researchers review and refine LLM-generated codes rather than accept generated themes.
Three concrete implications:
- Audit substitutions, not just accuracy. A code can be defensible and still have replaced the participant's vocabulary. Diff the AI's language against the transcript's language on any identity-related term.
- Treat theme-level output as a draft. Agreement degrades exactly where the analytic contribution lives. If a tool cannot show you the codes and quotes beneath a theme, you cannot check the step that fails most.
- Raise the bar for sensitive corpora. With vulnerable participants, the cost of flattening a claim is borne by someone other than the researcher.
This is why Qualitati's ThemeLens keeps every generated theme traceable to its underlying codes and verbatim quotes — the review step the paper argues for only works if the evidence chain is visible. For the data-collection side, an AI Interviewer that preserves full verbatim transcripts gives that audit something to check against.
FAQ
Does this study mean AI thematic analysis is invalid?
No. It found above-chance agreement with human analysis and no fabricated content, and its authors recommend LLM use for coding and clustering. The finding is that agreement falls as abstraction rises, so theme-level output needs human review.
What is adjusted mutual information (AMI), and is 0.23 good?
AMI measures how much two clusterings of the same items agree, corrected for chance, where 0 is chance and 1 is identical. The reported 0.231 is statistically above chance but far from equivalence — consistent with partial overlap rather than replication.
Would a different model have done better?
The study tested one model and does not answer this. But the divergences described — paraphrase, generalisation, softened terminology — reflect general tendencies of instruction-tuned language models rather than a single model's quirk, so treat model choice as unlikely to remove the need for review.
What should I do differently on a sensitive project?
Use AI for first-pass coding, keep every theme traceable to quotes, review any term the model substituted for a participant's own, and have a human researcher own the final theme set.
Primary source: Markelius, A., Dogan, F. I., Bailey, J., & Gunes, H. (2026). Evaluating Human and LLM-Generated Thematic Analysis in HRI for Vulnerable Populations: A Comparative and Ethical Analysis. 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), Kitakyushu, Japan. arXiv:2608.21420.
This is an independent editorial summary of third-party research; it is not affiliated with or endorsed by the authors.
Last updated: 26 August 2026