{"slug":"expertise-adaptive-ai-interviewer-local-llm-2026","title":"Expertise-Adaptive AI Interviewers: What a 2026 Study Found","description":"Can an AI interviewer adapt to each participant's expertise? A 2026 study of 246 interviews with a local LLM found 78.9% profiling agreement and kappa 0.80.","keywords":"expertise-adaptive AI interviewer, AI-moderated interviews, local LLM interviewer, adaptive follow-up questions, AI qualitative interviews, Llama 3.2 interviewer","date":"2026-10-09","author":"Qualitati Research Team","category":"AI User Research","readTime":"7 min read","content":"<div class=\"short-answer\"><p><strong>Short answer:</strong> An expertise-adaptive AI interviewer estimates how much each participant knows from their answers and adjusts the difficulty of its next question to match. In a 2026 study of 246 interviews, a small local LLM matched participants' self-reported expertise 78.9% of the time and scaled question complexity in line with it.</p></div>\n\n<p>Most AI interviewers ask everyone the same depth of question. A novice gets a question pitched at an expert and gives a thin answer; an expert gets a basic question and loses interest. A new preprint by Aisvarya Adeseye, Jouni Isoaho, Adeyemi Adeseye, Seppo Virtanen and Mohammad Tahir tests a different design: an interviewer that profiles each participant's expertise after every answer and pitches the next question accordingly. This post summarises what they built, what the numbers show, and what it means if you run AI-moderated interviews.</p>\n\n<h2>What did the study test?</h2>\n<p>The study tested whether a prompt-driven interviewer running on a small, locally hosted model can adapt question depth to each participant without losing the thread of the conversation. According to <a href=\"https://arxiv.org/abs/2610.11651\" target=\"_blank\" rel=\"noopener\">Adeseye et al. (2026)</a>, the system ran on Llama 3.2 3B, kept local so participant responses never went to an external LLM service.</p>\n<p>The interview topic was how employees use LLMs at work. Each interview had 16 participant-facing questions spread across five research areas: awareness and knowledge of LLMs, organizational applications, skills and training, data privacy and security, and organizational guidelines. Each area carried a priority (high, medium or low) and a question budget. Participants were adults with workplace experience, recruited through university and professional networks, and interviewed online in English.</p>\n\n<h2>How does an expertise-adaptive AI interviewer work?</h2>\n<p>It works as a pipeline of five prompt-driven modules that share one persistent record of the interview. The same model plays every role:</p>\n<ol>\n<li><strong>Generate system prompt.</strong> Turns the researcher's interview configuration into one reusable system prompt covering the interviewer's role, purpose, ethical guidelines, knowledge-handling rules, behavioural consistency and output format.</li>\n<li><strong>Generate initial question.</strong> Asks one accessible, open-ended, low-complexity opener from a high-priority area. It calibrates, but does not decide, the participant's level.</li>\n<li><strong>Profile expertise.</strong> After every answer, classifies the participant as Novice, Basic Knowledge, Advanced Knowledge or Expert, using four factors: terminology and specificity, conceptual correctness, reasoning depth, and application and evaluation.</li>\n<li><strong>Generate the next turn.</strong> Writes an acknowledgement of the answer, a transition to the next topic, and a next question whose complexity follows the current expertise level.</li>\n<li><strong>Validate uniqueness.</strong> Compares each candidate question with everything already asked and either accepts it or sends it back for regeneration.</li>\n</ol>\n<p>The \"evidence-traceable\" part refers to that shared interview-state record. It stores previous questions, answers, research-area coverage, expertise estimates and the uniqueness checker's decisions with their justifications, so a researcher can later see why each question was asked.</p>\n\n<h2>How accurate was the expertise profiling?</h2>\n<p>Fairly accurate. Adeseye et al. compared the model's expertise label with each participant's self-reported level, collected independently. The model agreed exactly in 194 of 246 interviews (78.9%) and was within one adjacent level in 98.4%. The linearly weighted Cohen's kappa was 0.80, and only four participants were misclassified by more than one level.</p>\n<p>The self-reported mix was reasonably balanced: 52 Novice (21.1%), 73 Basic Knowledge (29.7%), 76 Advanced Knowledge (30.9%) and 45 Expert (18.3%), so the result is not driven by one dominant group.</p>\n\n<h2>Did question difficulty actually change with expertise?</h2>\n<p>Yes. Coders rated the complexity of all 3,690 follow-up questions on a 1–4 scale (89.1% exact agreement between coders, weighted kappa 0.86). Average complexity rose steadily with the profiled level, with a Spearman correlation of .79 (p &lt; .001):</p>\n<table>\n<thead><tr><th>Profiled expertise level</th><th>Mean question complexity (1–4)</th><th>SD</th></tr></thead>\n<tbody>\n<tr><td>Novice</td><td>1.42</td><td>0.55</td></tr>\n<tr><td>Basic Knowledge</td><td>2.08</td><td>0.60</td></tr>\n<tr><td>Advanced Knowledge</td><td>2.91</td><td>0.63</td></tr>\n<tr><td>Expert</td><td>3.56</td><td>0.51</td></tr>\n</tbody>\n</table>\n<p><em>Source: Adeseye et al. (2026), arXiv:2610.11651.</em></p>\n\n<h2>Does a uniqueness check stop repetitive questions?</h2>\n<p>It catches a small but real share. The checker accepted 3,469 of 3,690 candidate questions (94.0%) on the first pass and sent 221 (6.0%) back as too similar to something already asked. Of those, 207 passed after one regeneration and 14 needed two or more. Repetition is one of the most common complaints about AI interviewers, so a cheap, explicit check like this is worth copying.</p>\n\n<h2>How did participants rate the interviews?</h2>\n<p>Highly. On 5-point scales, participants rated question relevance and coherence at 4.41, engagement at 4.32 and overall satisfaction at 4.38. These are self-reports without a fixed-sequence comparison group, so they show the adaptive interviewer was well received, not that it beats a non-adaptive one.</p>\n\n<h2>What are the limitations?</h2>\n<p>The authors name two. Only one model was tested, so it is unknown whether other local models behave the same way. And self-reported expertise is a reference point, not ground truth: people over- and under-rate their own knowledge, so 78.9% agreement measures agreement with self-perception. Beyond that, the study used one topic and one language, did not report latency, and gave no quantitative measure of transition quality. Researchers should also ask what adaptation does to comparability: if novices and experts receive different questions, cross-group comparisons need care at the analysis stage.</p>\n\n<h2>What this means for researchers</h2>\n<ul>\n<li><strong>Small local models are now workable interviewers.</strong> A 3B-parameter model handled profiling, question generation and repetition checks, which matters for studies where data cannot leave the institution.</li>\n<li><strong>Split the interviewer into roles.</strong> Separating profiling, generation and validation made each step inspectable. A single \"be a good interviewer\" prompt gives you nothing to audit.</li>\n<li><strong>Log the reasoning, not just the transcript.</strong> Keeping why each question was asked supports the audit trail that reviewers increasingly expect from AI-assisted qualitative work.</li>\n<li><strong>Pilot with known profiles.</strong> Before fielding, run a few pilot interviews with people whose expertise you already know and check whether the interviewer's pitch matches.</li>\n</ul>\n<p>If you want to try adaptive follow-up probing on your own discussion guide, the <a href=\"https://qualitati.com/ai-interviewer\">QualiTaTi AI Interviewer</a> runs text and voice interviews with researcher-defined topics and follow-ups, and <a href=\"https://qualitati.com/themelens\">ThemeLens</a> handles the thematic analysis of the transcripts afterwards.</p>\n\n<h2>FAQ</h2>\n<h3>What is an expertise-adaptive AI interviewer?</h3>\n<p>It is an AI interviewer that infers how much a participant knows from their answers and adjusts the complexity of later questions, so novices are not overwhelmed and experts are pushed further.</p>\n<h3>Which model did the study use?</h3>\n<p>Llama 3.2 3B, running locally, for all five modules. The authors list testing other local models as future work.</p>\n<h3>Is self-reported expertise a reliable benchmark?</h3>\n<p>Only partly. The authors acknowledge that self-assessment can over- or underestimate knowledge, so the 78.9% figure is agreement with participants' own view of their expertise.</p>\n<h3>Where can I read the paper?</h3>\n<p>The preprint, \"Evidence-Traceable Dynamic Interviewer Architecture for Expertise-Adaptive Qualitative Interviews Using Local LLMs\", is on <a href=\"https://arxiv.org/abs/2610.11651\" target=\"_blank\" rel=\"noopener\">arXiv (2610.11651)</a>. The authors note it is accepted for a Springer Nature book, <em>AI in Education: Pedagogy, Ethics, and Society</em>.</p>\n\n<p><em>Last updated: October 9, 2026. This is an independent editorial summary of third-party research; QualiTaTi is not affiliated with the authors.</em></p>\n","related":[{"slug":"synthetic-data-accuracy-digital-twins-forgotten-question-2026","title":"Synthetic Data Accuracy: Why Digital Twin Claims Disagree","description":"Why do digital twin accuracy claims range from 95% to chance? A 2026 synthetic data study explains the metrics and tests a screen for answerable questions.","category":"AI User Research","date":"2026-10-08","readTime":"7 min read","author":"Qualitati Research Team"},{"slug":"synthetic-participants-first-click-testing-gpt-2026","title":"Can GPT Predict Where Users Click? A 2026 First-Click Study","description":"Can synthetic participants replace UX testing? A 2026 study of 3,431 users found GPT's first-click predictions diverged from real behavior in 53% of tasks.","category":"AI User Research","date":"2026-10-07","readTime":"7 min read","author":"Qualitati Research Team"},{"slug":"custom-gpt-ai-interviewer-field-report-2026","title":"Can a Custom GPT Run Research Interviews? A 2026 Field Report","description":"Can a custom GPT run research interviews? A 2026 field report with 66 participants found high acceptance but generic probing and unverifiable AI summaries.","category":"AI User Research","date":"2026-10-02","readTime":"7 min read","author":"Qualitati Research Team"}]}