Can AI Do Thematic Analysis on Focus Groups? 2026 Study
Qualitati Research Team · 2026-06-03 · 7 min read
Last updated: June 3, 2026
Short answer
Yes, but only for part of the job. A blinded 2026 study in PLOS Digital Health found that large language models matched or slightly beat human analysts when applying a predefined codebook to a focus-group transcript (deductive coding), with a strict hallucination rate of just 1.2%. On open-ended theme generation (inductive coding) the models were more variable and weaker at interpreting tone, relationships, and latent meaning. The verdict: LLMs can augment thematic analysis but still need human verification.
What did the study test?
Researchers ran the first blinded head-to-head comparison of LLMs and human analysts on focus-group data in a healthcare setting. According to Hill et al. (2026), the team analyzed a single ~2-hour focus group of 7 participants with lived experience of long-term health conditions — a 12,172-word transcript about an AI-informed digital health proposal, recorded in July 2025.
Three AI approaches were pitted against two blinded human researchers:
- ChatGPT-5 (general-purpose, via API)
- Claude 4 Sonnet (general-purpose, via API)
- QualiGPT (a purpose-built LLM qualitative-coding application)
Both groups did the work twice: once deductively (applying a fixed framework of 10 codes grouped into 6 themes) and once inductively (generating themes from scratch — 14 codes across 5 themes).
How accurate is AI thematic analysis?
On deductive coding, the AI models were statistically non-inferior to humans — and two of them edged ahead. According to Hill et al. (2026):
- Human analysts reached 92.7% agreement (95% CI 91.6–93.9); LLMs reached 93.5% (95% CI 92.5–94.5).
- Cohen's κ was identical at 0.34 for both groups; Gwet's AC1 was 0.92 (human) and 0.93 (LLM).
- All LLMs cleared the non-inferiority bar (p < 0.0001); ChatGPT-5 and Claude 4 Sonnet were statistically superior to humans (p = 0.048 and p = 0.039).
- LLM specificity was high at 0.98, though sensitivity was lower for rare or implicit codes in both groups.
The low Cohen's κ (0.34) alongside very high raw agreement is a known artifact of imbalanced data — most codes were absent in most segments (code prevalence was only 7.8%), which deflates kappa. That is exactly why the authors also report Gwet's AC1, which is robust to prevalence and showed strong concordance.
Deductive vs inductive: where AI breaks down
The headline takeaway is that AI is reliable for structured coding and shaky for interpretive work. The contrast is stark by task.
| Dimension | Deductive coding | Inductive theme generation |
| AI vs human result | Non-inferior; 2 of 3 models superior | Only ChatGPT-5 reached non-inferiority (p = 0.043) |
| Strongest AI performance | Privacy & data concerns; consent & transparency; personalization | Descriptive themes: personalization needs, access barriers, data security |
| Weakest AI performance | Patient expertise; collaborative involvement in design | Inferential domains: peer support, care fragmentation |
| Human advantage | Detecting infrequent / implicit codes | Latent meaning, relational and affective nuance |
Per-theme Cohen's κ ranged widely, from −0.05 to 0.54. AI did best on concrete, system-level themes and worst on themes requiring interpersonal reasoning — empathy, trust, and the relational texture of patient narratives. As the authors put it, models "excelled at abstracting technical/system-level implications" but lagged humans on "latent meaning and relational conceptualization."
Do LLMs hallucinate quotes in qualitative analysis?
Less than you might fear, but not zero — and how you count matters. According to Hill et al. (2026):
- Strict hallucination rate (claims with no supporting evidence): 1.2% (SD 2.1%) — and two of the three models hallucinated nothing at all.
- Expanded rate (including segments attributed to the wrong code): 8.6% (SD 5.1%).
- Comprehensive error rate (adding partial matches): 12.4% (SD 5.1%).
The practical lesson: fabricated quotes are rare, but misattributed or partially-correct evidence is common enough that every AI-surfaced quote needs to be traced back to the transcript before it lands in a report.
What this means for researchers
The study's recommendations map cleanly onto a defensible workflow:
- Use AI for high-volume descriptive coding and for mapping text to a predefined framework — that is where it is now non-inferior to humans.
- Keep humans in charge of deeply interpretive analysis — latent themes, emotional nuance, and relational meaning.
- Verify every quote, report your error rate, and document exactly how the AI output fed into the final analysis.
This is the design philosophy behind a human-in-the-loop thematic analysis workflow: let the model do the first pass and the volume work, then have a researcher validate, relabel, and anchor every theme in real participant evidence. Qualitati's ThemeLens follows the same map-reduce pattern across many transcripts while keeping quotes traceable to source — and if you are running group sessions, see our guide to AI-moderated focus groups.
Limitations to keep in mind
This was a single focus group (n = 7 participants, one transcript) in one healthcare topic, so the absolute numbers will not generalize to every dataset. The deductive framework was also imbalanced, which complicates kappa interpretation. Treat the direction of the findings — strong on structured coding, weak on interpretation, low but non-zero hallucination — as more durable than any single percentage.
FAQ
Can AI replace human coders for thematic analysis?
Not for interpretive work. The 2026 PLOS Digital Health study found LLMs non-inferior to humans for deductive coding but weaker and more variable for inductive theme generation, especially on themes requiring interpersonal reasoning. The authors recommend AI as augmentation with mandatory human verification.
How often do LLMs hallucinate in thematic analysis?
In this study the strict hallucination rate (fully unsupported claims) was 1.2%, with two of three models at zero. But when misattributed and partial-match errors are included, the comprehensive error rate rose to 12.4% — which is why quote verification is essential.
Which is more reliable for AI: deductive or inductive coding?
Deductive. Applying a predefined codebook is where the models matched or beat humans. Open-ended theme generation produced more variable results, and only one model reached non-inferiority on that task.
Which models were tested?
ChatGPT-5, Claude 4 Sonnet, and QualiGPT (a dedicated qualitative-coding app), accessed via API in October 2025, compared against two blinded human researchers.
Primary source: Hill C, Dahil A, Simpson G, Hardisty D, Keast J, Pinn CK, Dambha-Miller H. "Large language models for thematic analysis in healthcare research: A blinded mixed-methods comparison with human analysts." PLOS Digital Health, April 3, 2026. Read the paper.
This article is an independent editorial summary of third-party research. Figures and quotations are drawn from the cited paper; we did not conduct this study.