Do AI Survey Probes Improve Data Quality? A 1,800-Person Test
Qualitati Research Team · 2026-07-11 · 7 min read
Do AI survey probes improve data quality? A 2025 randomized experiment with 1,800 respondents found that AI chatbot probing produced longer, more detailed open-ended answers and coded them into pre-set categories better than chance without any fine-tuning — but at a slight cost to respondent experience and with false positives inflated by acquiescence bias. AI probing helps, with guardrails.
What did the study test?
The study built a framework for "AI-assisted conversational interviewing" and tested it in a live web survey. According to Barari et al. (2025), 1,800 participants recruited from a non-probability online panel were randomly assigned to AI chatbots powered by large language models. The chatbots did two jobs at once: they dynamically probed respondents for elaboration on open-ended questions, and they coded those answers into researcher-defined categories in real time — a task usually done by humans after fielding.
Respondents were split across a control condition and two probing treatments: Treatment 1 asked confirmation-style probes ("Just to confirm, did you mean…?"), and Treatment 2 asked elaboration or relevance probes that pushed for more detail. This design lets us separate the effect of asking more from the effect of the AI simply being present.
How accurate is AI live coding of survey answers?
Reasonably accurate — better than random, but not human-grade on every question. According to Barari et al. (2025), the chatbots "performed better than random guesses" on live coding across most questions without survey-specific fine-tuning. The clear exception was news-source questions, where both accuracy and recall fell below 70%. For context, the authors note prior LLM evaluations have reported 91–94% classification accuracy on open-ended survey responses, so live, un-tuned coding sits below that ceiling.
The catch is how accuracy was measured. When respondents themselves confirmed the AI's coding, agreement looked high. But when trained human coders re-checked the same responses, precision was on average 20% lower than the respondents' own confirmations. That gap is the fingerprint of acquiescence bias: people tend to say "yes" to a confirmation probe even when the category is slightly off.
Does AI probing get better data — or just more of it?
Both, with a trade-off. Open-ended responses under probing were "more detailed and informative," per Barari et al. (2025). But richer data came at a cost to the respondent's experience, and probing raised friction: answer changes and back-tracking to earlier items were roughly twice as common in the confirmation-probing condition as in the control (88 vs. 43 instances), and interview duration was longer for both treatments — the elaboration-probing chats lasted about twice as long on average.
AI-assisted conversational interviewing: benefits vs. costs
| Dimension | What the study found | Implication for researchers |
| Answer richness | More detailed, informative open-ended responses | Better raw material for thematic analysis |
| Live coding accuracy | Above chance without fine-tuning; <70% on news sources | Treat AI codes as a draft, not a verdict |
| False positives | Inflated by acquiescence on confirmation probes | Avoid leading "did you mean…?" phrasings |
| Human vs. self-agreement | Coder precision ~20% below respondent confirmation | Validate a sample against human coders |
| Respondent experience | Slightly worse; longer interviews, more back-tracking | Cap probe depth; probe only where it pays |
What this means for researchers
The headline for practitioners: AI probing is a genuine data-quality lever, not a gimmick — but it is not free. Three takeaways follow directly from the evidence.
- Probe selectively. Elaboration probes buy the richest answers but cost the most time. Reserve them for the two or three questions where depth actually matters, rather than probing everything.
- Watch your probe wording. Confirmation probes ("did you mean X?") invite yes-saying and inflate false positives. Neutral, open elaboration probes reduce that bias.
- Keep a human in the loop on codes. The 20% precision gap between respondent confirmation and human coders is a reminder to validate a sample of AI codes before trusting the full set.
These lessons map onto how modern platforms are built. Qualitati's conversational surveys use AI-driven follow-up questions and branching so open-ended answers get probed in the moment, and its work on AI follow-up questions reflects the same finding — good probing deepens data, poor probing just adds burden. For turning those richer transcripts into themes, ThemeLens keeps a researcher in the analysis loop rather than treating a first-pass AI code as final.
How this fits the wider evidence
This study sits alongside a growing body of work on where AI helps and hurts in qualitative and survey research. It echoes findings that AI can code open-ended data competently but drifts on ambiguous categories, and that agreement metrics can flatter AI if you measure them the wrong way. For a related look at whether AI models agree with each other on themes, see our summary of multi-LLM thematic analysis.
FAQ
Can AI chatbots replace human survey coders?
Not entirely. In this study the AI coded better than chance without fine-tuning, but human coders disagreed with roughly a fifth of the confirmations respondents accepted. AI is a strong first pass that still needs human validation on a sample.
Does adding AI probes annoy respondents?
Slightly. Probing lengthened interviews and increased back-tracking and answer changes, and overall experience dipped a little. The effect was manageable, but it argues for probing only where the extra detail is worth the burden.
Why do AI-coded answers get false positives?
Mostly acquiescence bias. When the chatbot asks "did you mean X?", respondents lean toward "yes" even when X is not quite right, so categories get confirmed that a neutral human coder would reject.
How big was the study?
1,800 participants, randomly assigned across a control condition and two AI-probing conditions in a web survey experiment, per Barari et al. (2025).
Last updated: July 11, 2026.
This article is an independent editorial summary of third-party research (Barari, S., Angbazo, J., Wang, N., Christian, L. M., Dean, E., Slowinski, Z., & Sepulvado, B. (2025). "AI-Assisted Conversational Interviewing: Effects on Data Quality and Respondent Experience," arXiv:2504.13908). Qualitati is not affiliated with the authors, and all figures are drawn from the cited preprint.