Voice Answers in Surveys: Richer Data, More Nonresponse
Qualitati Research Team · 2026-09-11 · 8 min read
Short answer: Voice answers to open-ended survey questions are usually two to three times longer than typed answers and take less time to give, but far more people skip them. In a 2024 randomized smartphone experiment with 1,001 German respondents, 39% skipped an oral probe versus 8% for a written one. Offer voice as an option for depth, not as the only way to answer.
Last updated: September 11, 2026
Open-ended questions are where surveys get their "why." They are also where respondents give up: a thumb-typed answer on a phone is often a handful of words. Letting people speak instead is the obvious fix, and survey methodologists have now tested it in enough randomized experiments to say what it does and does not buy you. This guide summarizes that evidence on voice answers in surveys, adds what 2026 research shows about transcribing them automatically, and gives a decision matrix for when to use them.
Key takeaways
- Voice answers are longer. Across studies, spoken answers run roughly two to three times the length of typed ones.
- Voice answers are skipped more. Oral probes drew four to five times more nonresponse than written probes in one 2024 experiment.
- Longer is not automatically richer. The same experiment found both modes produced similar themes.
- Automatic transcription is the new weak link. A 2026 comparison of three speech recognition tools found each failed in a different way.
- Best practice is choice, not replacement. Keep text available, and audit transcripts before coding them.
What are voice answers in surveys?
Voice answers are responses to open-ended survey questions that participants record by speaking, usually into a smartphone, instead of typing. The survey stores the audio, and the answer is later transcribed (manually or with automatic speech recognition) and coded like any other open text. Dictation, where the phone keyboard converts speech to text on the fly, is a related but different option: the researcher only ever sees text.
What does the research show?
The evidence base comes mostly from randomized experiments in online panels in Germany and Spain. The pattern is consistent.
Answers get longer and faster
In a smartphone experiment on sensitive topics, Höhne, Gavras and Claassen (2024) found voice answers were more than twice as long as text answers on all four sensitive open questions, while requiring shorter response times. In a web-probing experiment, Lenzner, Höhne and Gavras (2024) reported about three times as many words in oral answers: 43.82 versus 14.18 words on the first probe, and 41.19 versus 12.08 on the second.
Many more people do not answer
The same web-probing study (N = 1,001, Forsa Omninet panel) found roughly 39% of the oral group did not answer the first probe, compared with 8% of the written group. Earlier work cited by Höhne and colleagues reported dropout of about 45% for voice answers against about 13% for text. Revilla, Couper, Bosch and Asensio (2020) found that voice recording led to substantially higher nonresponse than typing, and dictation to slightly higher nonresponse. Willingness is also uneven across people: Lenzner and Höhne (2022) studied who is willing to use voice inputs and why, which matters because selective nonresponse is a bias, not just a smaller sample.
Extra words do not always mean extra insight
Lenzner, Höhne and Gavras found that both modes resulted in similar themes, with oral answers mentioning slightly more themes only on the first probe. Their overall verdict was a mixed picture with no clear winner. Spoken answers carry more hesitation, repetition and context, which is valuable for interpretation but does not guarantee new codes.
Can AI transcribe voice survey answers reliably?
Not without checking. Revilla, Ochoa, Höhne and Couper (2026) transcribed 859 Spanish voice answers with three tools:
| Tool | Main strength | Main failure |
| Google Cloud Speech-to-Text | 94.2% of transcripts rated very clear | Failed to transcribe 26% of audio files, especially long ones |
| OpenAI Whisper | Produced a transcript for 99.8% of cases | At least one added word in 28.5% and a missing word in 19.2%; 79.6% valid |
| Vosk | 90% coverage | Weak punctuation; only 10.7% rated very clear |
The authors also compared human coding with GPT-4o coding of the transcripts. Agreement was weak to moderate, with Cohen's kappa between 0.00 and 0.63 depending on the quality dimension. The practical lesson: the transcription step silently shapes your data, and the failure modes differ (dropped files versus invented words). Pick a tool knowingly and spot-check audio against text.
The Voice Answer Decision Matrix
This matrix is Qualitati's own synthesis of the studies above. Use it to decide, question by question, whether voice is worth its nonresponse cost.
| Situation | Recommendation | Why |
| Narrative question ("Tell us about the last time…") | Offer voice alongside text | Length and context matter most here |
| Short factual or list question | Text only | Voice adds nonresponse without adding content |
| Key outcome question you must have from everyone | Text default, voice optional | Nonresponse on voice can be four to five times higher |
| Respondents with low literacy or typing difficulty | Offer voice prominently | Removes a barrier to answering at all |
| Answers will be coded automatically at scale | Voice only with a transcript audit | ASR tools drop files or add words |
| You need tone, pace or emotion | Voice, analyzed as audio | Transcripts discard acoustic signal |
Voice-answer readiness checklist
- Text input stays available on every voice-enabled question.
- The consent screen says audio is recorded, stored and transcribed.
- A pretest checks the record button and microphone permission on iOS and Android.
- You compare who answered by voice versus text before pooling results.
- A sample of transcripts (for example 10%) is checked against the audio.
- Your coding plan treats hesitation and repetition consistently across modes.
When not to use voice answers
- When full coverage on a question matters more than depth.
- When respondents answer in public or shared spaces, such as commuters or workplace panels.
- When you cannot store audio under your ethics approval or data policy.
- When the language or accent mix is poorly supported by your transcription tool.
Where Qualitati fits
Qualitati is an AI user research platform for product, UX and customer insights teams. Its conversational surveys ask AI-driven follow-up questions and can embed an AI-moderated interview question that runs in text or real-time voice, so a respondent who wants to talk can, while the rest of the survey stays structured. For spoken data, AI-moderated interviews run in voice or text in 10 languages, Voice Analytics extracts acoustic features such as pitch, loudness variability and speech rate, and ThemeLens synthesizes themes across up to 100 transcripts with participant-anchored quotes. The same caution applies inside any platform: review transcripts and codes before reporting.
Limitations of this evidence
Most experiments come from German and Spanish online panels and a small group of research teams, so results may differ in other countries, languages and populations. Nonresponse rates also depend heavily on interface design and on how the voice option is introduced. Treat the numbers above as direction, not as a forecast for your study, and pretest.
FAQ
Are voice answers better than typed answers in surveys?
They are longer and faster to give, but they are skipped far more often and do not reliably produce more themes. They are better for narrative depth, not for coverage.
Why do respondents skip voice questions?
Studies point to privacy concerns, being in a public place, discomfort with recording and technical hurdles such as microphone permissions. Willingness also varies across respondent groups.
Should I use dictation or voice recording?
Dictation keeps the answer as text and showed only slightly higher nonresponse than typing in Revilla and colleagues' 2020 study; recording kept richer audio but had substantially higher nonresponse.
Which speech recognition tool is best for survey answers?
No tool won outright in the 2026 JSSAM comparison. Google produced the clearest text but skipped many files; Whisper covered almost everything but added words. Audit a sample either way.
Can an LLM code voice answers instead of humans?
It can assist, but in the 2026 study GPT-4o agreed with human coders only weakly to moderately. Keep a human-reviewed codebook and check agreement before relying on it.
Bottom line
Voice answers in surveys trade coverage for depth. Offer them on narrative questions, keep text available, audit transcription, and check who chose each mode before pooling. To try spoken and typed answers side by side, start free with 30 credits or view transparent pricing.
This article is an independent editorial summary of published research. Methodology recommendations should be reviewed against your own study design and ethics requirements.