What Are the Ethical Risks of AI-Assisted Interviewing?
Qualitati Research Team · 2026-07-16 · 8 min read
AI-assisted interviewing means an LLM suggests follow-up questions to a human interviewer during a live semi-structured interview. A 2026 study by Zhang and colleagues put 17 interviewers through exactly that setup and found the risks are not mainly about question quality — they are about respect, accountability, and who is exposed when the AI is listening.
What did the study test?
The researchers ran a Wizard-of-Oz simulation of an AI follow-up assistant rather than deploying a real one. According to Zhang, Liu, Guan, Cai, and Carroll (2026), each session had three roles: the study participant acted as lead interviewer, one researcher played the interviewee ("Oz"), and a second researcher ("Wizard") requested one to three candidate follow-up questions from GPT-4o after each response, then selectively relayed or edited them. The lead interviewer heard the suggestions as if they came from a co-interviewer.
This design matters. It preserves the thing that most AI-interviewing pilots quietly remove: a human deciding, in the moment, whether a machine-generated question should be asked at all. The paper calls this "LLM-in-the-loop" — the model proposes, a person disposes.
The sample was 17 interviewers (10 women, 7 men, aged 21–35) with mixed qualitative-methods experience: 10 novice, 5 intermediate (1–3 peer-reviewed qualitative papers), and 2 advanced (3 or more papers). The authors are explicit that this skews toward novice and intermediate researchers.
One methodological choice is worth flagging because it shapes everything that follows. Participants first evaluated the follow-up questions without being told they came from an LLM, and only afterwards were told and asked to reflect on ethics and responsibility. So the study captures both the naive reaction and the informed one.
What were the five concerns?
Participants raised five interlocking concerns. They are worth reading as a set, because each one is a different answer to "who bears the cost when this goes wrong?"
| Concern | What it looks like in practice | Who absorbs the harm |
| Harmful or discriminatory language | The model produces "very bizarre remarks," discriminatory phrasing, or questions that cross a boundary the interviewer would not have crossed | The interviewee |
| Undermined respect | The interviewer's attention splits between the person and the screen; nonverbal cues get missed | The interviewee |
| Participation inequality | Technical barriers exclude some people — the paper describes this as knowledge-based discrimination | Would-be participants |
| Unclear responsibility | When the AI generates something harmful, no one is clearly accountable — model vendor, researcher, or institution | Nobody, which is the problem |
| Privacy and compliance | An AI that listens, records, or transcribes sensitive content creates disclosure and regulatory exposure | The interviewee and the institution |
Why is the "respect" finding the interesting one?
Because it is a cost that no amount of better prompting fixes. The first concern — bad language from the model — is a tractable engineering problem; you can filter, constrain, and review. The second is structural. An interviewer reading AI suggestions is, by definition, not fully looking at the person in front of them. The study's participants experienced that divided attention as a form of disrespect toward the interviewee, and as a loss of the nonverbal signal that tells a skilled interviewer when to push and when to stop.
That trade-off is the real design question for anyone building or buying this technology. A follow-up assistant that improves probing depth while degrading rapport may produce transcripts that look richer and interviews that felt worse. Neither an automated quality metric nor a word count will catch that.
What about the accountability gap?
The fourth concern is the one most likely to matter to research ethics boards. In a conventional interview, if a question causes harm, the interviewer asked it and the interviewer answers for it. In the LLM-in-the-loop setup, that chain has a new link — and the study's participants could not settle where responsibility lands when a relayed AI question does damage.
This is not a hypothetical. It determines what goes in your consent form, what your ethics board approves, and what happens when a participant complains. The paper's response is a mapping from each risk category to mitigation strategies: bias-aware prompts, mandatory human oversight, accessible interfaces, explicit accountability frameworks, and privacy-by-design with regular audits. None of these is exotic. All of them require someone to have decided, in advance, who is on the hook.
What this means for researchers
Three things follow if you are considering AI assistance in live interviews.
- Keep a human in the relay path, and say so. The study's whole design assumed a person could edit or drop any suggestion. A tool that fires AI questions directly at participants is testing a different, riskier thing than what this paper evaluated.
- Decide your disclosure policy before fielding, not after. The study deliberately separated pre-disclosure and post-disclosure reflection, which suggests the authors expected knowing to change the judgment. Assume your participants' reactions will change too.
- Treat transcription and recording as the compliance surface. The fifth concern is where most institutional risk actually lives — not in the cleverness of the questions, but in what the system hears and stores.
If you are running fully AI-moderated interviews rather than AI-assisted ones, these concerns shift rather than disappear. The divided-attention problem goes away when there is no human interviewer to distract; the accountability and privacy questions get sharper, because the model is now the only one asking. Our AI Interviewer is built around explicit probing rules and reviewable transcripts for that reason, and ThemeLens keeps the analysis step auditable rather than opaque. But the honest reading of this paper is that tooling choices are downstream of governance choices, not a substitute for them.
Frequently asked questions
Does AI-assisted interviewing improve probing depth?
This study did not measure probing depth as an outcome. Its stated motivation is that interviewer cognitive load and limited domain familiarity can constrain probing, but the findings reported here are about ethical concerns rather than data quality. Treat depth gains as an open question.
Is a Wizard-of-Oz study a fair test of a real AI tool?
It is a fair test of the human-oversight version. The Wizard could edit and withhold suggestions, which approximates a deployment where a researcher reviews AI output before use. It does not tell you what happens in a fully automated system.
How generalizable are the findings?
Cautiously. Seventeen participants, mostly novice or intermediate qualitative researchers, aged 21 to 35 — the authors themselves note the skew. The concerns are a useful agenda for discussion, not a population estimate.
Should I disclose AI assistance to interviewees?
The paper's privacy and responsibility findings point strongly toward yes, and most institutional review processes will expect it. The study's own design — testing reactions before and after disclosure — implies disclosure is consequential enough to plan for deliberately.
Source
Zhang, H., Liu, Y., Guan, X., Cai, J., & Carroll, J. M. (2026). Ethics and Social Responsibility in AI-Assisted Interviewing: An LLM-in-the-Loop Study of AI-Generated Follow-Up Questions. arXiv:2606.30980. Accepted to CHIWORK '26.
Last updated: July 16, 2026
This is an independent editorial summary of third-party research. Qualitati is not affiliated with the authors, and any interpretation beyond the paper's own claims is our own. Readers should consult the original paper for its full argument and limitations.