Do AI Follow-Up Questions Help Qualitative Interviews?
Qualitati Research Team · 2026-06-02 · 6 min read
Do AI-generated follow-up questions improve qualitative interviews? A 2026 study found they can deepen exploration and reduce interviewer workload, but only when a human stays in control of timing and relevance. AI worked best as a "springboard" suggesting probes — not as an autonomous interviewer firing questions on its own schedule.
What did the study test?
Researchers tested whether AI-generated follow-up questions (AGQs) help human interviewers run better semi-structured interviews. In the paper "Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews" (Zhang, Liu, Guan, Cai & Carroll, 2026), the team used an AI-driven Wizard-of-Oz design: a researcher posed as a co-interviewer while GPT-4o generated real-time follow-up questions during live interviews. Participants acted as the lead interviewer and decided whether to use each suggested probe.
Who took part?
The study involved 17 participants (ages 21–35; 10 female, 7 male), all with prior qualitative research experience. According to Zhang et al. (2026), 9 had prior qualitative analysis or interview experience and 8 had peer-reviewed publications using qualitative methods — an experienced sample whose judgments about interviewing quality carry weight.
How were the AI questions delivered?
The study varied when and how often AI probes appeared, across three insertion patterns:
- Occasional single questions — one AGQ inserted now and then.
- Periodic insertions — an AGQ after several participant-led questions.
- Multiple consecutive questions — a burst of AI probes in sequence.
This let the researchers observe how pacing — not just content — shaped the interview experience.
What did interviewers think of the AI probes?
Most found them genuinely useful for going deeper. Participants described AGQs as "connectors" or "springboards" that pushed topics further and acted as "buffers" giving them thinking space. According to Zhang et al. (2026), 14 of 17 participants said the AI maintained good thematic relevance to the interview. Most also reported reduced cognitive load — relief from the mental work of formulating the next question on the fly — and a few (P3, P8, P9, P12) noted emotional support and added confidence.
Where did the AI fall short?
Timing was the biggest weakness. A majority of participants (P5, P13, P14, P15, P16) flagged poor timing as the core problem: early or badly placed probes disrupted a planned interview flow, and not knowing when the next AI contribution would land created uncertainty. Participants also raised the risk of offensive or culturally insensitive questions and worried about comfort for vulnerable interviewees. The consistent conclusion: human gatekeeping was strongly preferred — interviewers wanted to approve, reshape, or drop each suggestion rather than have it inserted automatically.
AI interviewer vs. AI co-pilot: what's the difference?
| Dimension | AI as autonomous interviewer | AI as human-controlled co-pilot (this study) |
| Who decides timing | The model | The human interviewer |
| Main benefit | Scale, consistency | Deeper probing, lower interviewer load |
| Main risk | Disruptive or insensitive questions | Slower pace; depends on interviewer skill |
| Relevance control | Hard to guarantee | Human filters every probe |
| Best fit | High-volume, structured topics | Sensitive, exploratory, depth-first studies |
What this means for researchers
For teams running qualitative interviews, the practical takeaway is to treat AI as a suggestion engine the interviewer commands, not a hands-off moderator. Surface candidate follow-ups, but keep a human in the loop to control pacing and screen for sensitivity. This mirrors how human-in-the-loop tooling is built into platforms like Qualitati's Interview Assistant, where AI proposes probes a researcher can accept or edit live. Once interviews are done, the same human-oversight principle applies to analysis — AI can draft codes and themes in ThemeLens, but researchers validate them.
How strong is the evidence?
This is early, qualitative evidence. With 17 participants, a single model (GPT-4o), simulated interview scenarios, and self-reported perceptions rather than measured interview-quality outcomes, the findings are directional, not definitive. They map the experience of using AI probes well, but they don't yet quantify whether AGQs produce richer data than a skilled human alone.
FAQ
Can AI generate good follow-up questions in interviews? Yes — in this 2026 study, 14 of 17 experienced researchers found AI follow-ups thematically relevant and useful for deepening topics, provided a human chose when to use them.
Should AI run interviews on its own? The study's participants strongly preferred human gatekeeping. Poorly timed or insensitive AI questions were the main risk, so a human-controlled co-pilot model is safer for sensitive or exploratory work.
What's the biggest weakness of AI interview probes? Timing. Most participants said badly placed AI questions disrupted the interview's flow more than the content of the questions did.
Primary source: Zhang, H., Liu, Y., Guan, X., Cai, J., & Carroll, J. M. (2026). Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews. arXiv:2509.12709.
Last updated: June 2, 2026. This article is an independent editorial summary of third-party research and is not affiliated with or endorsed by the authors.