Can a Custom GPT Run Research Interviews? A 2026 Field Report
Qualitati Research Team · 2026-10-02 · 7 min read
Short answer: A custom GPT can run short, self-administered research interviews that participants accept. In a 2026 experience report with 66 software practitioners, 90.9% rated the experience positively and 89.4% would take part again. But only 19.7% said the follow-up questions helped them say more, and the study does not show the data match human interviews.
Many researchers have wondered whether they could skip building an interview platform and simply configure a ChatGPT "GPT" to interview their participants. A new preprint tests exactly that do-it-yourself setup in two real studies. According to Gheyi, Albuquerque, Ribeiro and Perkusich (2026), the workflow is workable and well received for 10–15 minute interviews, but it leaves gaps in probing depth, privacy, and data fidelity that any research team should plan for before copying it.
What did the study test?
The study examined whether a customized ChatGPT agent (a "MyGPT") could conduct semi-structured interviews without a researcher present. The paper is AI-Conducted Interviews in Empirical Software Engineering: An Experience Report (arXiv, July 16, 2026), by researchers at the Federal University of Campina Grande and the Federal University of Alagoas in Brazil.
The team built two GPTs, one per study. The first asked software professionals about refactoring practices; the second asked professionals and students how they use generative AI in Scrum activities such as backlog refinement and user story writing. The setup worked like this:
- Prompt design. Each GPT had a shared core (role, tone, one question at a time, ask for concrete examples, avoid judgmental language, avoid collecting confidential information) plus a topic-specific question list. The instruction field was capped at roughly 8,000 characters, which forced concise prompts.
- Pilot. Researchers tested access through a shared link (including from free accounts), voice mode, protocol adherence, and language switching in three languages.
- Fieldwork. Participants opened the link with their own ChatGPT account, chose a language, and spoke with the interviewer by voice, with no researcher present.
- Output. At the end, participants asked the GPT for a structured synthesis (profile, practices, risks, "possible codes") and pasted it into a Google Form along with a short experience questionnaire.
Over three days the team analyzed 66 submissions, 33 per study. All participants lived in Brazil, and 65 of the 66 artifacts were written mainly in Portuguese.
How did participants rate being interviewed by AI?
Participants rated the experience highly on convenience and clarity, and more cautiously on naturalness and on replacing a human. According to Gheyi et al. (2026):
| Questionnaire item | Positive responses (n = 66) |
| Pace was appropriate | 97.0% (64/66) |
| Questions were clear | 95.5% (63/66) |
| Positive overall experience | 90.9% (60/66) |
| Felt comfortable answering | 90.9% (60/66) |
| Would participate again | 89.4% (59/66) |
| Able to express experiences and opinions | 87.9% (58/66) |
| Interview felt natural | 81.8% (54/66) |
| Absence of a human researcher was positive | 43.9% (29/66) |
The pattern is telling. Logistics scored highest: 65.2% valued answering at their own pace and 47.0% said they felt less pressure than with a human researcher. But the absence of a human was rated positively by fewer than half, with 45.5% neutral. On a comparative item, 42.4% chose the midpoint, 37.9% leaned toward a human interviewer, and 19.7% leaned toward AI. The authors caution that this item's wording was ambiguous, so it signals a broad orientation, not a firm preference.
Where did the AI interviewer fall short?
The weak spot was adaptive probing. While 65.2% said the AI kept the interview focused, only 19.7% said its follow-up questions helped them give more detailed answers. The most frequently selected weaknesses were:
- Questions too generic: 28.8%
- Limited sensitivity to their answers: 28.8%
- Insufficient depth: 27.3%
- Privacy concerns: 22.7%
- Missed human interaction: 22.7%
- Repetitive questions: 19.7%
The authors' own conclusion is that a shared prompt buys consistency, and consistency should not be confused with good probing. Voice also added noise: in one case a cough was taken as an answer, and some participants had to repeat requests to change language.
Can you trust an AI-generated interview summary?
Not on its own. This is the most important methodological point in the paper. The researchers never collected full transcripts; they received only the synthesis each participant chose to paste in. Their audit found that 92.4% (61/66) of submissions had the expected structure, but three contained conversation excerpts instead, one held unstructured answers, one had no usable content, and two contained refactoring content although the participant had reported doing the Scrum interview.
Part of the problem was wording: the form asked for a "transcript" while the GPT produced a summary. More fundamentally, a summary can compress, reorder, or omit what a participant said, and the study could not check what was lost. The authors explicitly treat these artifacts as AI-mediated summaries, not data equivalent to transcripts, and they note that the GPT's "possible codes" are not validated researcher codes.
What this means for researchers
The study supports a narrow but useful claim: for short, focused, low-risk topics, a self-administered AI interview is acceptable to participants and can gather completed responses quickly. If you try it, the paper's lessons translate into a practical checklist:
- Keep the full transcript. A participant-pasted summary is not data you can audit. Use a setup where the researcher, not the participant, holds the complete record.
- Write participant instructions as carefully as the prompt. Name the exact artifact you expect, and check every submission on arrival so you can ask for a resubmission.
- Control the platform. When participants use their own accounts, you cannot fix the model, the interface, or data retention. Tell participants plainly what the platform may store.
- Log the language and the model for each interview, so variation can be reported.
- Do not rely on AI interviews alone for sensitive or emotional topics or for discussions of screens, code or documents.
- Ask participants how it felt. A short post-interview questionnaire shows how the method was experienced, not only whether it produced data.
A purpose-built tool addresses several of these gaps at once. Qualitati's AI Interviewer stores the full text or voice transcript for the researcher, runs under the project owner's configuration rather than each participant's account, and lets you set a probing-depth level plus conditional follow-up probes in the interview guide. Its transcripts can then go straight into ThemeLens for thematic analysis where every theme links back to quotes, instead of to an AI summary.
Limitations of the study
- It is a preprint and an exploratory experience report, not an experiment, and there is no human-interviewer comparison.
- Only completers were analyzed; invitations and drop-outs were not tracked, so no completion rate can be reported.
- The sample was 66 people from one country across two technical topics, with interviews of about 10–15 minutes.
- The artifact audit was done by one researcher, and results are tied to OpenAI's GPT platform at the time of the study.
FAQ
Can I use a custom GPT as an AI interviewer for my study?
For short, low-risk interviews it can work: in this 2026 study 89.4% of 66 participants would do it again. Plan for generic follow-ups, participant-side privacy settings you cannot control, and the lack of a researcher-held transcript.
Do participants prefer AI interviewers to human ones?
Not clearly. Most were comfortable, and 47.0% felt less pressure, but on the comparative item more leaned toward a human interviewer (37.9%) than toward AI (19.7%), with 42.4% neutral.
Is an AI-generated interview summary good enough for qualitative analysis?
No. Summaries can omit or reorganize what participants said, and in this study 5 of 66 submissions did not contain the expected summary at all. Analyze full transcripts and treat summaries as navigation aids.
What kinds of studies suit self-administered AI interviews?
Short, focused studies on low-sensitivity topics where scheduling is the main bottleneck. The authors advise against using them as the only method for sensitive, emotional, or artifact-heavy discussions.
Last updated: October 2, 2026
This article is an independent editorial summary of third-party research by the Qualitati Research Team. It is not affiliated with or endorsed by the study's authors. Consult the original paper for full methods and results.