Should You Translate Interviews Before AI Analysis?
Qualitati Research Team · 2026-09-04 · 9 min read
Short answer. Translate interviews before analysis only when your analytic team cannot read the source language. Coding in the original language and translating the output — themes, definitions, and selected quotes — preserves the meaning that matters and reduces the number of lossy steps from two to one. Translate-first is a staffing decision disguised as a methods decision.
The decision nobody documents
Every multilingual project reaches the same fork. You have 40 interviews in Portuguese, Japanese, and Arabic. Your analysis will be written in English. Somewhere between the recording and the report, a translation happens. The only real question is where — and the answer changes what your findings are.
Two orders are available:
- Translate-first. Transcribe in the source language, translate every transcript to English, then code the English.
- Source-language-first. Transcribe and code in the source language, then translate the codebook, themes, and the quotes you actually quote.
Most teams pick translate-first without deliberating, because it makes the corpus look uniform and lets one analyst handle everything. That convenience is real. So is the cost, and the cost is rarely written down. A 2022 review in the International Journal of Qualitative Methods found that reporting of the translation process in qualitative health research is routinely neglected — readers frequently cannot tell what was translated, by whom, or at what stage (Yunus et al., 2022). If it is not reported, it was not scrutinized.
Key takeaways
- Translate-first adds a lossy step before interpretation, so every downstream code inherits the translation's compressions.
- Source-language-first moves the lossy step after interpretation, where you can see what you are compressing and choose.
- Machine translation quality is not uniform across languages. A 2026 benchmark of 24 bidirectional language pairs found BLEU scores dropping 68.7% from legal to literary text, and near-zero scores on some low-resource pairs (arXiv:2510.07877). Interview talk is closer to the literary end than the legal one.
- LLM thematic analysis has been shown to work directly in at least one non-English language: a 2024 study ran inductive thematic analysis on Italian semi-structured interviews with Italian prompts and got themes resembling those produced independently by human researchers (De Paoli, 2024).
- Neither order removes the need for a bilingual human at the point where meaning is decided.
Why the order matters more than the quality
The instinct is to argue about translation accuracy. That is the wrong axis. Assume for a moment that your translation is excellent. It is still a reading — a set of resolved ambiguities. "Complicado" became "complicated" rather than "awkward." A Japanese participant's sentence-final hedge became a period. An Arabic diminutive of affection became a neutral noun.
Each of those is a small interpretive decision. In translate-first, they are all made before coding, by someone (or something) whose job was translation, not analysis, and who did not know which distinctions your research question cares about. The coder then codes the resolutions, not the ambiguities, and has no way to see that a choice was made.
In source-language-first, the same ambiguity is still there — but it reaches the person who is looking for it. They can code it as ambiguous. They can flag it. When it turns out to matter, they can write about it.
This is why the methodological literature has long favored analyzing in the language of collection where feasible. The AI question does not change the logic; it changes the feasibility.
What changed in 2026: feasibility, not principle
Source-language-first used to be gated on staffing. You needed an analyst fluent in every language in the corpus. For a study in six languages, that meant six analysts, or a compromise.
Two developments have moved the gate — and it is worth being precise about how far.
LLMs can code in languages they support. The Italian study above is a targeted existence proof: prompted in Italian, on Italian transcripts, a pre-trained model produced codes of usable quality. The paper's own framing is conditional — suitable for multilingual work "so long as the language is supported by the model used." That condition is doing real work.
Support is uneven, and predictably so. The 2026 translation benchmark is the clearest public evidence of the shape of that unevenness. Cross-language-family pairs underperform intra-family pairs; the gap narrows with model scale but does not close. Bias detected in the outputs concentrated in low-resource source languages, with cultural and sociocultural bias accounting for over 75% of flagged instances across 1,439 human-annotated pairs. That study measured translation, not coding — but coding in a language the model handles poorly is not obviously safer than translating out of it.
So the honest 2026 position is not "AI solved multilingual qualitative analysis." It is: for well-supported languages, source-language coding is now cheap enough to be the default; for poorly supported ones, you have traded one known risk for a less-studied one.
The analysis-language decision matrix
Score your project on the four rows below. This is a practitioner heuristic, not a validated instrument.
| Factor | Favors source-language-first | Favors translate-first |
| Research question | How people talk about it — framing, metaphor, hesitation, identity, emotion | What people report — features requested, prices named, steps taken, tasks failed |
| Language support | High-resource language, well covered by your model and ASR | Low-resource language where the model codes unreliably but a human translator is available |
| Team | At least one reviewer reads the source language | No reader of that language anywhere on the team, and none hireable |
| Output | Publication, regulatory file, anything where quotes carry evidentiary weight | Internal readout where directional signal is enough |
Three or four rows on the left: code in the source language and translate only what you quote. Three or four on the right: translate-first is defensible — document it and treat quotes as paraphrase. A split verdict usually means the third row is the binding constraint, and the fix is a reviewer, not a method.
The hybrid nobody names
There is a third order that gets used in practice and almost never described in method sections: code in the source language, translate the codes, reconcile across languages, then quote from the source with a translated gloss. It costs one extra reconciliation meeting and it is what most rigorous multilingual projects converge on. Give it a name in your write-up so a reviewer can evaluate it.
Source-language coding audit — six checks
If you code in the original language, whether by hand or with AI assistance, these are the six checks that catch the failures specific to this route.
- Prompt language matches data language. An English prompt over a Portuguese transcript quietly invites the model to think in English and back-translate internally. Match them, and say which you used.
- Code names live in one language, definitions in two. Pick a working language for code labels so the codebook stays sortable, but write each definition in both. Divergence between the two definitions is the earliest signal that the code means different things in different corpora.
- Cross-language reconciliation is a scheduled step. Two analysts coding two languages will produce two codebooks. Merging them is analysis, not admin, and it needs a slot on the calendar.
- A bilingual reviewer spot-checks 10% per language. Not the whole corpus. Enough to catch a systematically mis-applied code.
- Quotes are published in source plus translation. Both. A translated-only quote is unfalsifiable by your reader.
- Language-specific behavior is logged. If the model refused, hedged, or degraded in one language, that is a finding about your instrument and belongs in the limitations section.
Who this is for — and when not to bother
Use source-language-first if: your corpus spans languages, your question is about meaning rather than counts, at least one person on the team reads each language, or the output will be scrutinized externally.
Skip the debate if: every participant was interviewed in the analysis language already; or the study is a fast internal signal-check where directional findings are enough and nobody will quote a participant verbatim. Do not spend a week on translation architecture for a two-day readout.
Where Qualitati fits
Qualitati is an AI user research platform for product, UX, and customer insights teams, covering AI-moderated interviews and focus groups, conversational surveys, and AI-assisted qualitative analysis.
Qualitati supports research in 10 languages — English, Chinese, French, Norwegian, Dutch, German, Spanish, Portuguese, Japanese, and Arabic — across the interview itself and the analysis that follows. Practically, that means an interview conducted in Japanese can be analyzed as Japanese: ThemeLens runs thematic analysis as a map-reduce pipeline across up to 100 transcripts at once, mapping codes to research questions and synthesizing themes with participant-anchored quotes, and QDA Workspace supports inductive and deductive coding and codebook generation. The source-language route does not require you to hand-build a pipeline per language.
What the platform does not do is make the reconciliation decision for you. Merging a Japanese codebook with a German one is interpretive work, and the six checks above still apply. Nor does language support mean equal support: the caution from the 2026 benchmark holds inside any tool, ours included, and a language your model handles poorly will code poorly.
Pricing is published per credit, with a free tier of 30 credits on signup and no credit card required.
Limitations and honest caveats
- The direct evidence is thin. We are not aware of a published head-to-head trial coding the same corpus both ways and comparing the resulting themes. That study should exist and does not, as far as we can find. Everything above is reasoning from adjacent evidence.
- The Italian study is one language, one dataset. It establishes that non-English LLM thematic analysis can work. It does not establish that it works in your language, on your data, at your standard.
- The translation benchmark measured translation. Extending its language-family and low-resource findings to coding quality is an inference, not a measurement.
- "Supported language" has no agreed definition. Vendors, ours included, publish language lists without a per-language quality floor. Treat any such list as a starting point for your own pilot, not a guarantee.
- Cost is not modeled here. Translating 40 transcripts and translating 30 themes are very different invoices, and which is cheaper depends on rates we cannot generalize.
Human-review note: the decision matrix and six-check audit are practitioner frameworks synthesized from cross-language qualitative methodology, not validated instruments. Have a methodologist review them before they enter a preregistration or an ethics submission.
Bottom line
Whether you translate interviews before analysis should follow from who can read the source language, not from what makes the spreadsheet tidy. Translation is interpretation; put it after coding, where an analyst can see the choice, rather than before, where it arrives already decided. Where the language is well supported, coding in the original is now the cheaper and more defensible default — and where it is not, say so in your limitations rather than assuming the model handled it.
FAQ
Should I translate interviews before qualitative analysis?
Only if no one on the analysis team reads the source language. Translating before coding means every code is applied to someone else's resolved reading of the data. Coding in the source language and translating the themes and selected quotes afterward keeps the interpretive decision with the analyst.
Can AI do thematic analysis in languages other than English?
Yes, for languages the model supports well. A 2024 study ran inductive thematic analysis on Italian semi-structured interviews using Italian prompts and found the resulting themes resembled those produced independently by human researchers. The paper's own caveat — that this holds "so long as the language is supported by the model used" — is the operative condition.
Is machine translation good enough for research transcripts?
It depends heavily on the language pair and the register. A 2026 benchmark across 24 bidirectional pairs found translation quality dropped 68.7% from legal to literary text and approached zero on some low-resource pairs. Interview speech — with hedges, idiom, and emotional register — behaves more like literary text than legal text.
What is the biggest risk of the translate-first approach?
Invisible compression. The translation resolves ambiguities that your research question may have been about, and the coder never learns a choice was made. A translated transcript reads as fluent and complete, which is precisely why the loss goes unnoticed.
Do I still need a bilingual researcher if I use AI?
Yes. AI changes how much bilingual labor you need, not whether you need any. At minimum, one reviewer per language spot-checking a sample of coded segments, and verifying every quote that appears in the final output.
How should I report the translation step in a paper?
State what was translated (transcripts, codes, or quotes only), at what stage, by whom or by what system, and whether any verification was performed. Reporting of the translation process is routinely omitted in published qualitative work, so being explicit is a low-cost credibility gain.
Try it on your own data
Qualitati offers AI-moderated interviews, AI-moderated focus groups, conversational surveys, and AI-assisted thematic analysis in 10 languages, with published per-credit usage rates. Start free with 30 credits — no credit card required — or view transparent pricing. If you are evaluating a move off legacy QDA software, see our NVivo alternative guide and MAXQDA vs ATLAS.ti vs NVivo comparison.
Related reading: multilingual qualitative research with AI, why back-translation fails for qualitative research, and code-switching in multilingual interviews. On the broader question of what AI-assisted coding can and cannot be trusted with, see Schroeder et al. (2024) on the norms gap researchers report.
Last updated: September 4, 2026. This article is an independent editorial summary. Competitor and tool claims reflect publicly available information as of that date; where a detail is not published, we say so rather than estimate.