Can Open-Source LLMs Code Qualitative Interviews? 2026 Study
Qualitati Research Team · 2026-06-24 · 7 min read
Last updated: June 24, 2026
Short answer
Open-source LLMs can speed up qualitative coding, but they are not yet a substitute for a human coder. In a 2026 study in the International Journal of Qualitative Methods, Gemma2 and Llama3.1 coded 34 patient interviews — and roughly half of the AI-generated codes lacked context, were duplicative, or did not align with researcher themes. Deductive (framework-guided) coding clearly outperformed inductive coding, and the models occasionally surfaced codes researchers had missed.
What the study tested
The study asked a practical question: can freely available, open-source language models do the first pass of qualitative coding well enough to help a research team? According to Misra, Dahal, Kirk, Khan, Dogan, Chataut, and Gyawali (2026), in "Large Language Models in Qualitative Analysis: Comparing Traditional and Researcher-Interpreted Approaches", the researchers analyzed 34 semi-structured interviews with adults managing chronic disease in a 12-week diabetes lifestyle intervention.
They tested two open-source models — Gemma2 and Llama3.1 — under two conditions: a deductive approach (coding guided by a predefined framework) and an inductive approach (codes emerging from the data). Traditional, researcher-led qualitative analysis served as the human baseline, and the team then evaluated every machine-generated code for context, duplication, and alignment with the themes researchers had identified by hand.
How accurate was the AI coding?
Mixed — strong on volume, weak on relevance. The models produced thousands of codes, but a large share were noise. According to the authors, "nearly half of the LLM generated codes lacked sufficient context, were repetitive, and had poor alignment with the themes developed by researchers."
Two patterns stand out. First, the models over-generated: Llama3.1 produced 2,040 total codes under the deductive approach (1,042 unique), versus Gemma2's 1,666 (723 unique). Second, alignment with human themes was low across the board — the best condition (Gemma2 deductive) reached only 27.1% "good fit" with researcher themes.
Deductive vs inductive coding: the numbers
The single clearest finding is that framework-guided (deductive) coding beat open-ended (inductive) coding on nearly every metric. The table below summarizes the reported results from Misra et al. (2026).
| Metric | Gemma2 (deductive) | Llama3.1 (deductive) | Gemma2 (inductive) | Llama3.1 (inductive) |
| Total codes generated | 1,666 | 2,040 | 1,530 | 1,632 |
| Unique codes | 723 | 1,042 | 715 | 829 |
| Codes providing context | 45.6% | 43.7% | 38.9% | 22.9% |
| Good fit with researcher themes | 27.1% | 26.8% | 19.2% | 11.7% |
| Reviewer agreement (inter-rater) | 94% | 78% | 74% | 87% |
Duplication was a persistent problem: the authors report that 29.5%–36.2% of deductive codes and 16.3%–28.5% of inductive codes were duplicative. In other words, the more codes a model generated, the more a researcher had to spend time deduplicating and discarding.
Did the AI add anything human coders missed?
Yes — and this is the most encouraging result. Misra et al. (2026) note that the models "generated a few additional codes that were missed previously by the researchers, that could have allowed strengthening themes or developing additional subthemes." That points to a genuine, if modest, use case: LLMs as a bias-reducing second pass that flags overlooked patterns, not as the primary coder.
What this means for researchers
Treat open-source LLMs as an assistant on the first pass, not an autonomous coder. Three practical takeaways from the study:
- Use a framework. Deductive coding produced more contextual, better-aligned codes than inductive coding in every model tested. If you have an existing codebook, give it to the model.
- Budget for cleanup. With up to a third of codes duplicative and roughly half off-target, the human work shifts from generating codes to validating, merging, and pruning them.
- Keep a human in the loop. The models' low theme alignment means final themes still belong to the researcher. AI-assisted coding works best when every machine code is human-reviewed.
This matches how purpose-built tools approach the problem. Platforms like QDA Workspace support both inductive and deductive coding with researcher review built into the loop, and ThemeLens maps codes back to research questions and anchors every theme in participant quotes — directly addressing the "lacks context" failure mode the study identified.
Limitations of the study
This was a single dataset (34 interviews on one health topic) and two specific open-source models. Larger or proprietary models, different prompts, or a different domain could shift the numbers. The "good fit" and context judgments were also made by the research team, so they reflect one expert group's interpretation. The findings are best read as a realistic snapshot of mid-2025-era open-source LLM coding, not a ceiling on what AI coding can do.
FAQ
Can open-source LLMs replace human qualitative coders?
Not yet. In Misra et al. (2026), roughly half of AI-generated codes were off-target, duplicative, or lacked context, and the best theme alignment was 27.1%. They work as an assistant, not a replacement.
Is deductive or inductive AI coding better?
Deductive. Framework-guided coding outperformed inductive coding on context and theme alignment for both Gemma2 and Llama3.1 in this study.
Which open-source models were tested?
Gemma2 and Llama3.1, chosen to balance accessibility and performance.
What is the main risk of using LLMs for coding?
Volume without relevance: the models generated thousands of codes, many duplicative or weakly tied to the data, shifting researcher effort toward validation and cleanup.
The bottom line
Open-source LLMs can accelerate qualitative coding and occasionally catch patterns humans miss, but in this 2026 study they produced as much noise as signal. The reliable workflow is human-led, framework-guided, and AI-assisted — with a researcher reviewing every code. If you want that workflow built in, see how ThemeLens and QDA Workspace keep humans in control. Start free with 30 credits.
This article is an independent editorial summary of third-party research. It is not affiliated with or endorsed by the study's authors or publisher. Please consult the original paper for full methodology and results.