Quote Selection Bias in AI Thematic Analysis (2026)
Qualitati Research Team · 2026-08-20 · 10 min read
Last updated: August 20, 2026
Short answer
Quote selection bias in AI thematic analysis is the tendency of an AI pipeline to illustrate themes with quotes drawn disproportionately from certain participants, phrasings, or emotional registers. It matters because published evidence shows LLM coding errors are not random with respect to participant characteristics, so a theme can look well-supported while systematically under-representing part of your sample.
Key takeaways
- Two peer-reviewed studies show LLM-assisted qualitative coding shifts label distributions in non-random ways — by participant characteristics, and by anchoring human reviewers.
- Quote selection is where that bias becomes visible to your stakeholders, because quotes are what they read and remember.
- Prevalence counts and quote counts are different claims. Most AI outputs blur them.
- Use the Quote Evidence Audit and Voice Coverage Matrix below before a readout, not after.
- The fix is structural: verbatim-anchored quotes, per-participant coverage, and a human pass on the tail, not just the top themes.
What quote selection bias actually is
Every thematic analysis ends the same way: a theme name, a sentence of interpretation, and two or three quotes underneath it. Those quotes carry almost all of the persuasive weight in the room. Nobody in a stakeholder readout re-reads 40 transcripts; they read the quotes and form a mental image of “the user.”
Quote selection bias is what happens when the passages chosen to illustrate a theme are not representative of the passages that actually support it. The theme can be real and the quotes can be genuine, and the readout can still mislead — because the four people quoted are not the twenty-two people coded.
This is an old problem in qualitative research. Human analysts have always gravitated to articulate, vivid, emotionally legible speakers. What is new in 2026 is that AI pipelines reproduce that gravitational pull at scale, silently, and with the appearance of systematic coverage.
Three mechanisms behind it
- Fluency preference. Passages that are grammatical, self-contained and quotable are easier for a model to extract and rank as strong evidence. Participants who hedge, ramble, speak a second language or think out loud produce fewer clean candidates.
- Salience preference. Emotionally intense language is easier to match against a theme label like “frustration with onboarding” than a flat, procedural description of the same experience.
- Position effects. Long transcripts are chunked, and evidence concentrates in whichever chunks a summarization step retained. If your pipeline summarizes before it selects, quotes drift toward the parts of the interview the summarizer liked.
What the evidence says
Two peer-reviewed studies are worth reading in full before you trust an AI-generated evidence table.
Julian Ashwin, Aditya Chhabra and Vijayendra Rao, writing in Sociological Methods & Research (first published online May 27, 2025), analyzed 2,407 open-ended interview transcripts with Rohingya refugees and Bangladeshi hosts in Cox's Bazar — 32,463 question-answer pairs coded into 19 categories. Their central finding is that LLM coding errors were not random with respect to the characteristics of the interview subject: the models over-predicted codes, and prediction errors correlated with refugee status, gender and parental education. They report that bespoke supervised models trained on a subset of sociologist-coded transcripts outperformed the LLMs they tested on both accuracy and bias. (DOI: 10.1177/00491241251338246)
The practical translation: if error rates differ by participant group, so does the pool of passages your pipeline believes supports each theme. Quote selection inherits the bias in coding, one layer up.
The second study addresses the standard defense — “we keep a human in the loop.” Hope Schroeder, Jad Kabbara and Deb Roy, in a pre-registered experiment published in Findings of ACL 2025, gave 410 annotators over 7,000 subjective annotation tasks across three AI-assistance conditions, two models and two datasets. Showing LLM suggestions did not make annotators faster, but it did raise their self-reported confidence, and annotators took the suggestions strongly enough to significantly change the label distribution versus the no-assistance baseline. (Findings of ACL 2025)
The practical translation: a reviewer who is shown the AI's three suggested quotes and asked “do these support the theme?” is not an independent check. Reviewing an AI shortlist is a different task from selecting from the full pool, and it produces different answers.
Add the known hallucination risk in citation-style tasks: an AI-generated quote that is paraphrased rather than verbatim, or attributed to the wrong participant, is a serious problem in a qualitative report and is not visually distinguishable from a correct one. Any pipeline that cannot show you the exact source passage and its location is asking for trust it has not earned.
Prevalence claims vs. evidence claims
Most of the damage from quote selection bias is done by a single unstated slippage: treating “how many people said this” and “how many quotes we show” as the same number.
| Claim type | What it means | What it needs | Common failure |
| Prevalence | N of 24 participants expressed this | Per-participant coding, deduplicated | Counting coded segments instead of people, inflating a talkative participant into a trend |
| Evidence | Here is what it sounded like | Verbatim quote, participant ID, location | Paraphrase presented as quotation; no traceable source |
| Intensity | This mattered a lot to those who raised it | Explicit intensity criteria, applied consistently | Reading intensity off the vividness of the language, which is a speaker trait |
| Distribution | This concentrated in a subgroup | Coverage broken out by segment | Reported only when convenient; never checked for themes that flatter the hypothesis |
If your AI output gives you a theme, a paragraph and three quotes, it has given you an evidence claim. Everything else in that table is your responsibility to establish.
The Quote Evidence Audit
This is a Qualitati-original check. Run it on every theme in a deliverable, before the readout. Score each item pass or fail; a theme with two or more fails is not ready to present.
| # | Check | Passes when |
| 1 | Verbatim match | Every quote is found by exact string search in the source transcript, with participant ID and a timestamp or line reference. No paraphrase is formatted as a quotation. |
| 2 | Speaker spread | The quotes shown for the theme come from at least three participants, and no single participant supplies more than half of them. |
| 3 | Coverage ratio | You can state how many distinct participants the theme is coded in, not just how many segments — and that number appears next to the theme. |
| 4 | Tail sample | A human has read at least three supporting passages the pipeline ranked lowest, not only the ones it ranked highest. They still fit the theme. |
| 5 | Register range | Not every quote is emotionally intense. At least one flat, procedural or ambivalent passage supporting the theme is included or explicitly considered. |
| 6 | Segment check | Coverage is broken out by at least one participant characteristic that matters for the study (role, tenure, language, market, accessibility need), and any concentration is stated in the theme description. |
| 7 | Disconfirming pass | Someone searched the corpus for passages that contradict the theme, and the deliverable says what was found — including “nothing substantive.” |
Checks 1–3 are mechanical and can be automated. Checks 4–7 are the human work that AI assistance makes easier to skip and more important to keep. See also our guides on negative case analysis and the AI qualitative analysis audit trail.
The Voice Coverage Matrix
The audit above works theme by theme. The matrix works across the study, and it catches what a per-theme audit cannot see: a participant who is coded everywhere but quoted nowhere, or a participant who supplies a quarter of all quotes in the report.
Build a table with one row per participant and these columns:
- Themes coded in — how many of your final themes this participant contributes evidence to.
- Quotes used — how many of their passages appear in the deliverable.
- Quote share — their quotes as a percentage of all quotes in the report.
- Words in transcript — a rough proxy for how much they said.
- Segment — the characteristic you care about for this study.
Then read it for three patterns:
- Silent contributors. Coded in three or more themes, zero quotes used. These are usually your least fluent speakers, and they are being coded but not heard.
- Over-representation. Any participant above roughly a fifth of total quotes in a study of ten or more people. Sometimes justified; always worth a sentence in the methods note.
- Segment skew. A segment that is a third of your sample and a tenth of your quotes. Fix the quotes, or say so in the report.
None of these thresholds are laws. They are trip-wires; the point is that somebody looked.
Where Qualitati fits
Qualitati is an AI user research platform for product, UX and customer insights teams. It runs AI-moderated interviews and focus groups in text and voice, conversational surveys, and AI-assisted analysis — and it is built so the evidence behind a theme stays traceable.
- ThemeLens runs a map-reduce thematic analysis across up to 100 transcripts at once, mapping codes to your research questions and synthesizing themes with participant-anchored quotes. Because it maps codes at the transcript level before reducing to themes, the per-participant coverage needed for checks 2 and 3 is a property of the pipeline rather than something you reconstruct afterward.
- QDA Workspace supports inductive and deductive coding, codebook generation and theme visualization, so you can inspect the coded segments a theme rests on — including the ones no quote was drawn from.
- Multilingual research in ten languages (English, Chinese, French, Norwegian, Dutch, German, Spanish, Portuguese, Japanese and Arabic) matters here specifically: fluency-driven quote selection punishes second-language participants hardest when everyone is forced into one interview language.
- Voice Analytics extracts acoustic features from interview audio — pitch, loudness variability, speech rate, voice quality — giving an intensity signal that is not simply who used the most colorful words.
What no platform can do for you is check 7. Searching your own corpus for evidence against your own theme remains a human decision.
Limitations and trade-offs
- Balanced quoting is not automatically better quoting. If one participant genuinely had the most insightful account of a phenomenon, quoting them twice is correct. The matrix flags concentration for explanation, not for automatic correction.
- The cited studies are not about your study. The Ashwin et al. corpus is development-research interviews in Bangladesh; the Schroeder et al. experiment used crowdworkers on subjective annotation tasks. Both establish that non-random bias occurs in LLM-assisted coding. Neither gives you an error rate for your domain, your language or your model version. Treat them as a reason to audit, not as an effect size to quote at stakeholders.
- Auditing costs time. A full Voice Coverage Matrix on a 40-participant study is real work. For low-stakes exploratory rounds, checks 1–3 alone catch most of the damage.
- Human review is not a free fix. That is the whole point of the Schroeder et al. result. If you want an independent check, have the reviewer select from the full pool of coded passages rather than approve the AI's shortlist.
- Methodology note for sensitive work. If your findings will inform a decision affecting a specific population — accessibility, healthcare, financial inclusion — have a second qualified researcher review the coverage breakdown before publication.
Who this is for, and when not to use this approach
Who this is for: UX researchers and insights leads presenting AI-assisted thematic analysis to stakeholders; research ops teams setting quality standards for a team using AI QDA tools; academics using LLM assistance who need a defensible methods paragraph.
When not to use this approach: if your study has fewer than about eight participants, the matrix produces noise rather than signal — read every transcript instead. If you are running a fully manual analysis with two human coders and formal inter-rater reliability, this audit is largely redundant with what you already do. And if the deliverable is a raw quote bank rather than a synthesized set of themes, there is no selection step to audit yet.
Bottom line
AI-assisted analysis did not create quote selection bias; it industrialized it and hid it behind a clean output format. The published evidence as of August 2026 is clear on one point: LLM coding errors distribute unevenly across participants, and human reviewers who see AI suggestions inherit the skew rather than correcting it. Auditing quotes is the cheapest place to catch this, because quotes are the part of your report people actually read.
FAQ
What is quote selection bias in qualitative research?
It is the systematic over-representation of certain participants, phrasings or emotional registers among the quotes used to illustrate a theme. The theme may be valid and the quotes genuine while the evidence shown still misrepresents who supported it.
Does AI make quote selection bias worse?
It makes it faster and less visible. Ashwin, Chhabra and Rao (Sociological Methods & Research, 2025) found LLM coding errors correlated with participant characteristics including refugee status, gender and parental education, so the pool of passages a pipeline treats as supporting evidence is itself unevenly distributed.
Doesn't a human reviewer solve this?
Not by default. In a pre-registered experiment published in Findings of ACL 2025 with 410 annotators and over 7,000 annotations, showing LLM suggestions significantly changed the resulting label distribution and raised annotator confidence without making them faster. Review the full candidate pool, not the AI's shortlist.
How do I check that an AI-generated quote is real?
Exact string search against the source transcript, and require a participant ID plus a timestamp or line reference for every quote in the deliverable. Paraphrase should never be typeset as quotation. If your tool cannot show the source passage in context, that is the finding.
How many participants should a theme be quoted from?
There is no universal rule. As a working trip-wire, show quotes from at least three participants per theme, keep any one participant under half the quotes for that theme, and always state the number of distinct participants the theme is coded in alongside it.
Can I automate this audit?
Checks 1 to 3 — verbatim match, speaker spread and coverage ratio — are mechanical and should be automated or built into your platform. Checks 4 to 7 — tail sampling, register range, segment breakout and the disconfirming pass — are judgment calls and should stay human.
Conclusion
Guarding against quote selection bias costs an hour and protects the credibility of everything downstream of your readout. Run the Quote Evidence Audit on each theme, build the Voice Coverage Matrix once per study, and write one honest sentence in your methods note about who you quoted and why.
Start free with 30 credits — no credit card required. Create an account and run an AI-moderated interview, focus group, conversational survey or thematic analysis project with participant-anchored evidence you can trace. See transparent pricing, or read our related guides on writing up thematic analysis findings, auditing AI qualitative coding and why agreement is not quality. More about Qualitati.
Sources: Ashwin, Chhabra & Rao, “Using Large Language Models for Qualitative Analysis can Introduce Serious Bias,” Sociological Methods & Research (2025); Schroeder, Kabbara & Roy, “Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks,” Findings of ACL 2025; arXiv preprint 2507.15821. All accessed August 20, 2026. This is an independent editorial summary; it is not affiliated with or endorsed by the authors or their institutions.