What Researchers Say About LLM Thematic Analysis (2025)
Qualitati Research Team · 2026-06-02 · 8 min read
Last updated: June 2, 2026
Short answer
Can large language models do thematic analysis on their own? Experienced researchers say no — not yet. In a 2025 reflective-workshop study of 25 software-engineering researchers, participants saw real gains in efficiency and scalability from LLM-assisted thematic analysis but flagged serious risks around bias, contextual loss, and reproducibility, concluding that LLMs can support interpretive analysis but not substitute for it (Ornelas et al., 2025).
What did the study test?
The study asked how experienced researchers conceptualize the opportunities, risks, and methodological implications of folding LLMs into thematic analysis. According to Ornelas et al. (2025), the authors ran a structured reflective workshop with 25 researchers from the International Software Engineering Research Network (ISERN). Rather than benchmarking a model against a gold-standard codebook, the design captured expert judgment: participants used color-coded canvases to document perceived opportunities, limitations, and recommendations across three stages of analysis — LLM-assisted coding, theme generation, and review.
That framing matters. Most coverage of "AI thematic analysis" measures accuracy on a fixed dataset. This study instead asks the people who do thematic analysis for a living where they would and would not trust the tool — which is closer to the decision a working researcher actually faces.
What are the opportunities researchers recognized?
Participants did not dismiss LLMs. The headline benefits were efficiency and scalability: the ability to move through large volumes of qualitative data faster than manual coding allows, and to support early, exploratory passes over a corpus. In a field where hand-coding hundreds of interview transcripts can take weeks, that speed is a genuine draw.
The optimism, though, was conditional. The same researchers who valued the speed were explicit that it only pays off when paired with disciplined human judgment — a theme that runs through every recommendation below.
What are the limitations of LLM thematic analysis?
According to Ornelas et al. (2025), participants converged on three recurring risks:
| Risk | What it means in practice | Why it matters for rigor |
| Bias | The model's training and prompting can tilt which codes and themes surface | Themes may reflect the model's priors, not the data |
| Contextual loss | Nuance, tone, and situated meaning get flattened when text is chunked and summarized | Interpretive depth is the whole point of qualitative work |
| Reproducibility | The same prompt and data can yield different themes across runs and model versions | Findings become hard to audit or replicate |
Reproducibility is the quietest but most corrosive of the three. Thematic analysis already lives with the charge that it is subjective; adding a non-deterministic model that shifts behavior between versions makes "we got these themes" harder to defend unless the process is documented carefully.
What did researchers recommend?
The study's recommendations centered on two ideas. First, prompting literacy: researchers need to understand how prompt wording, examples, and framing shape what the model returns, rather than treating the LLM as a black box that "finds the themes." Second, continuous human oversight: a human stays in the loop at every stage — checking codes, interrogating generated themes, and owning the interpretation.
The overall conclusion, in the authors' framing, is that LLMs as tools can support but not substitute interpretive analysis. That lines up with a broader pattern in 2025–2026 methods research: AI is most defensible as an accelerant for the mechanical parts of coding, with meaning-making left to the researcher.
A practical division of labor
- Let the model do the first pass. Use it to draft codes and cluster large volumes of text — the labor-intensive, low-judgment work.
- Keep theme generation human-led. Treat AI-proposed themes as candidates to challenge, not conclusions to accept.
- Document the prompts and model version. Reproducibility risk shrinks when your prompt, examples, and model are recorded alongside the findings.
- Validate on a sample. Hand-check a slice of the AI's codes against your own reading before trusting the full set.
What this means for researchers
If you run thematic analysis — in software engineering, UX, market research, or academic work — the takeaway is not "avoid LLMs" or "automate everything." It is to design a workflow where the model accelerates coding while a human owns interpretation and auditability. Tools that keep the researcher in control of codes and themes, rather than hiding the reasoning, fit this brief best. Qualitati's ThemeLens is built around exactly this human-in-the-loop model for AI thematic analysis, and pairs naturally with an AI Interviewer when you are also collecting the data. Whatever tool you use, the study's bar is a fair one: speed is welcome, but the themes are still yours to defend.
FAQ
Can LLMs replace researchers in thematic analysis?
No. According to Ornelas et al. (2025), experienced researchers concluded that LLMs can support but not substitute interpretive analysis, with human oversight needed at every stage.
What are the biggest risks of AI thematic analysis?
The study highlights three: bias in which themes surface, loss of context and nuance, and poor reproducibility across runs and model versions.
What is "prompting literacy"?
It is the researcher's understanding of how prompt wording, examples, and framing shape an LLM's output — treated in the study as a core skill for using these tools responsibly.
How was the study conducted?
It was a structured reflective workshop with 25 ISERN software-engineering researchers, who documented opportunities, limitations, and recommendations on color-coded canvases.
This article is an independent editorial summary of third-party research. Primary source: Ornelas, T., Araújo, A. A., Araújo, J., Araújo, M., Trinkenreich, B., & Kalinowski, M. (2025). LLM-Assisted Thematic Analysis: Opportunities, Limitations, and Recommendations. arXiv preprint.