Can AI Replace Human Thematic Analysis? A 2025 Study
Qualitati Research Team · 2026-06-30 · 7 min read
Last updated: June 30, 2026
Short answer
No — not on its own, but it changes the workflow. A 2025 proof-of-concept study found that GPT-4o reached thematic saturation 5-6 times faster than human coders and matched humans on inter-rater reliability for deductive coding (K = 0.84), but produced shallower inductive themes and shorter, single-code excerpts. The practical takeaway: use LLMs to accelerate analysis, keep humans for interpretive depth.
Key takeaways
- According to Parkington et al. (2025), an out-of-the-box GPT-4o model reached theme saturation at 15-20 transcripts, a knowledge-base-augmented version at 10-15, while human analysts needed 90-99.
- The out-of-the-box LLM showed strong inter-rater reliability against humans (K = 0.84) — but on deductive (theory-driven) coding, its strongest task.
- LLMs "lacked the depth of human analysis": they underperformed on inductive theme development and synthesis, and typically applied a single code per excerpt where humans applied several.
- The study's verdict: LLM thematic analysis is "more cost-effective" and can "transform qualitative analysis... when combined with human oversight" — a human-in-the-loop conclusion, not a replacement one.
What did the study test?
The researchers ran a head-to-head comparison of human and LLM-based thematic analysis on the same dataset: interview transcripts from n = 20 healthcare workers who took part in a stress-reduction trial. Human analysts coded the transcripts in Dedoose, a standard qualitative-data-analysis tool. The LLM side used OpenAI's GPT-4o in two configurations — a plain "out-of-the-box" model and a "knowledge-base" version primed with domain material — both driven by a structured prompt built on the RISEN framework (Role, Instructions, Steps, End-goal, Narrowing).
Because both sides analyzed identical data, the design isolates the effect of who (or what) is doing the coding, rather than confounding it with different source material — the right way to ask whether AI can stand in for a human coder.
How fast does AI reach saturation vs. humans?
The headline finding is about saturation — the point where additional transcripts stop surfacing new themes. The three approaches diverged sharply:
| Analyst | Saturation point (transcripts) | Relative speed |
| Knowledge-base GPT-4o | 10-15 | Fastest |
| Out-of-the-box GPT-4o | 15-20 | Fast |
| Human analysts (Dedoose) | 90-99 | Baseline |
That gap cuts two ways. The optimistic read: an LLM can map the thematic landscape of a study in a fraction of the transcripts, which is a real efficiency win for large datasets. The cautious read: reaching saturation faster is not the same as reaching it correctly. If the model stops finding new themes at transcript 15 because it is converging on surface-level patterns, early saturation is a symptom of shallowness, not a strength. The study's own depth findings support the cautious interpretation.
How accurate is the AI coding?
On reliability, the out-of-the-box GPT-4o earned a Cohen's kappa of 0.84 against human coders — conventionally read as "almost perfect" agreement. But the number deserves a caveat the study makes explicit: the LLM excelled at deductive coding (applying a pre-defined codebook) and underperformed at inductive development (building themes up from the data) and at theme synthesis.
That split matters because the two tasks demand different things. Deductive coding is closer to pattern-matching against fixed categories — exactly where current LLMs are strong. Inductive thematic analysis requires holding ambiguity, weighing context, and noticing what is not said. The study observed that humans produced longer excerpts carrying multiple codes, while the LLM tended toward shorter, single-code selections — a quantitative trace of the interpretive depth that automated coding still misses.
What this means for researchers
The result lines up with a growing 2025-2026 consensus: LLMs are a powerful assistant for thematic analysis, not a replacement for the researcher. A defensible workflow looks like this:
- Let the model do the first pass. Use an LLM to apply your codebook deductively and to surface candidate themes quickly across the corpus.
- Treat early saturation as a hypothesis, not a stop signal. Spot-check later transcripts by hand to confirm the model is not converging prematurely.
- Reserve humans for inductive work. Theme synthesis, edge cases, and the multi-code, context-heavy passages are where human judgment still pays off.
- Report your reliability. If you lean on AI coding, publish the human-vs-AI agreement for your own data rather than assuming a study's K transfers.
This is the model behind tools like ThemeLens, which runs a map-reduce thematic-analysis pipeline across transcripts and anchors every theme to participant quotes so a researcher can audit the synthesis rather than trust it blind. For codebook-driven deductive work — the LLM's strongest mode in this study — the QDA Workspace supports AI-assisted inductive and deductive coding with human review built in. For more on the human-AI division of labor, see our guide to inductive vs. deductive coding and LLM-human inter-rater reliability.
Limitations to keep in mind
This was a proof-of-concept study with n = 20 participants in a single mental-health context, using one model family (GPT-4o). The authors frame it as exploratory, and the saturation and reliability figures should be read as directional rather than definitive. Domain, prompt design, and model version all move these numbers. The honest summary is that the study demonstrates a pattern — fast, reliable deductive coding paired with weaker inductive depth — not a fixed performance benchmark.
FAQ
Can GPT-4o do thematic analysis? Yes, with caveats. In this 2025 study it reached saturation faster than humans and hit K = 0.84 inter-rater reliability on deductive coding, but it produced shallower inductive themes and single-code excerpts, so it works best with human oversight.
What is theme saturation? The point in qualitative analysis where new transcripts stop revealing new themes. In this study, GPT-4o reached it at 10-20 transcripts versus 90-99 for humans — though faster saturation may reflect shallower analysis, not greater efficiency.
Is AI thematic analysis reliable? For deductive coding against a fixed codebook, this study found "almost perfect" agreement (K = 0.84). For inductive theme development and synthesis, the LLM was weaker, so reliability depends heavily on the task and should be measured on your own data.
Should I use AI instead of a human coder? The study's conclusion is "AI plus human," not "AI instead of human." LLMs cut cost and time on the mechanical parts; humans still carry the interpretive depth.
This article is an independent editorial summary of third-party research. It synthesizes findings from Parkington et al. (2025), "Human vs. LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study" (arXiv:2507.08002), and is not affiliated with or endorsed by its authors.