Do LLMs Agree With Human Coders? A 2025 Reliability Study
Qualitati Research Team · 2026-05-30 · 7 min read
Last updated: May 30, 2026
Short answer
In a 2025 study, GPT-4o and GPT-4.5 reached substantial agreement with human coders (Cohen's kappa > 0.6) on three of four themes in a real qualitative dataset, but only moderate agreement on a domain-general theme. LLMs can match human inter-rater reliability on concrete codes — but still need human oversight on abstract constructs.
Where Qualitati fits
Qualitati's QDA Workspace and ThemeLens let AI propose and apply codes across transcripts while researchers review and validate every code — speed with human-in-the-loop rigor. View transparent pricing.
Full article available at qualitati.com/blog/llm-human-coder-inter-rater-reliability-2026.