How to Pretest Survey Questions With AI (2026)
Qualitati Research Team · 2026-09-08 · 9 min read
Short answer: AI survey pretesting means using an LLM to simulate cognitive interviews on draft questions before you field them. In the strongest published test to date, a guided simulated-cognitive-testing prompt caught 75% of deliberately embedded flaws in a 20-item questionnaire, in under an hour and for a few dollars (Sturgis, Roberts & Robinson, April 2026). Use it to screen; keep humans to confirm.
Key takeaways
- The best LLM configuration in the Survey Futures study detected 75% of known question problems with a small false-positive rate, at a cost of a few dollars per 20-item instrument (Sturgis, Roberts & Robinson, 2026).
- Simulated cognitive testing and expert review catch different problem types. Run both; neither is sufficient alone.
- Results were highly sensitive to prompt design and model choice — structured prompts that worked on open-weight models performed poorly on the proprietary model tested.
- Directive probes inflate agreement. In a 2025 study, respondents affirmed unlikely interpretations more than five times as often under directive probing than under open probing (Conrad et al., 2025). Scripted AI probes are structurally at risk of the same bias.
- The practical workflow is a funnel: LLM screen the whole instrument, then run real cognitive interviews on the flagged minority.
What is AI survey pretesting?
AI survey pretesting is the use of a large language model to evaluate draft survey questions before data collection, either by reviewing the wording against known problem taxonomies (simulated expert review) or by simulating a respondent who thinks aloud and answers probes about each question (simulated cognitive interviewing).
The traditional method it imitates is cognitive interviewing, formalized around Tourangeau's four-stage model of survey response: comprehension, retrieval, judgment, and response mapping. A participant answers a draft question while narrating their reasoning, or responds to targeted probes, and an analyst identifies where the question breaks down.
Cognitive interviewing works. It is also slow and expensive, which is why it is rationed. Sturgis and colleagues note that US academic organizations typically run a median minimum of around five respondents and a median maximum of around twenty per round — small samples that will reliably surface only the most prevalent problems. On a 60-item instrument, exhaustive human testing of every question is simply not on the table.
What the 2026 evidence actually says
The most direct evidence comes from Survey Futures Working Paper 15 (Patrick Sturgis and Thomas Robinson, LSE; Caroline Roberts, University of Lausanne), released April 2026. The design is a detection task rather than a demo: 20 questions with deliberately embedded problems, plus 7 items from the European Social Survey as a control set of well-designed questions. Three simulated-cognitive-testing prompt variants were compared against a simpler expert-review prompt.
Headline results, as reported by the authors:
- Detection. The best configuration — a guided cognitive testing variant — detected 75% of known problems with a small false-positive rate.
- Complementarity. Simulated cognitive testing and expert review identified different kinds of problems, mirroring the long-standing human finding (Presser & Blair, 1994) that the two methods do not substitute for each other.
- Fragility. Detection and false-positive rates were sensitive to prompt design and model choice. Structured prompts that worked well on open-weight models performed poorly on the proprietary model tested. Open-weight models achieved higher detection but much poorer discrimination between flawed and well-designed items — they flagged more, including things that were fine.
- Economics. A full evaluation of a 20-item questionnaire completed in under an hour for a few dollars.
Read the fragility finding carefully, because it is the one most likely to bite you. A pretesting method whose hit rate moves with prompt phrasing is not a measurement instrument yet — it is a screening heuristic. That distinction determines where it belongs in your workflow.
A second line of evidence points the same way. In "Exploring LLMs for Automated Generation and Adaptation of Questionnaires" (CUI '25; Adhikari, Hartland, Weber & Cannanure), a pipeline that generated, LLM-pretested, and culturally adapted questionnaires was evaluated with 238 US and 118 South African participants. Respondents rated LLM-pretested text as more specific and LLM-adapted questions as slightly clearer and less biased than the comparison versions. Useful, incremental, not transformative.
The Question Pretesting Triage Matrix
This is the Qualitati framework for deciding which method to spend on which item. The premise: LLM screening is nearly free, human cognitive interviews are not, and the scarce resource is respondent time — so spend it only where an error would be consequential and an LLM is unlikely to find it.
| Item stakes ↓ / LLM detectability → | High LLM detectability (wording, double-barrels, ambiguous scales, missing options) | Low LLM detectability (cultural referents, sensitive framing, domain jargon, presuppositions) |
High stakes (key DV, segmentation driver, pricing item, regulated content) |
Screen then confirm. LLM pass first, then 5–8 human cognitive interviews on the revised wording. The LLM saves the first revision cycle, not the last. |
Human first. Go straight to cognitive interviews with in-segment participants. Use the LLM only afterward, as a second pair of eyes on the revised item. |
Low stakes (descriptive, exploratory, firmographic) |
LLM only. Accept the residual risk. This is where the 75% detection rate is straightforwardly a win over the zero pretesting these items usually get. |
LLM plus expert review. Skip respondent time. Have one methodologist read the flagged items against a problem taxonomy. |
The matrix does one thing: it stops teams from treating "we ran it past an AI" as a completed pretest on the items where being wrong is expensive.
How to run an AI cognitive pretest: a five-step protocol
Step 1 — Build the problem taxonomy first
Before you prompt anything, write down what you are looking for. Working from a taxonomy rather than an open "find issues" instruction is exactly what distinguished the guided variant in the Survey Futures study from weaker prompts. A workable starting list, mapped to Tourangeau's stages:
- Comprehension: ambiguous terms, undefined jargon, double-barreled questions, vague quantifiers ("often", "regularly").
- Retrieval: reference periods that are too long, requests for counts nobody tracks.
- Judgment: presuppositions, leading framing, socially desirable answer directions.
- Response mapping: non-exhaustive or overlapping options, missing "not applicable", scale labels that do not match the stem.
Step 2 — Run two independent passes, not one
Pass A is simulated expert review: give the model the item and the taxonomy, and ask for flagged problems with the taxonomy category and a rewrite. Pass B is simulated cognitive interviewing: give the model a specific respondent persona, ask it to answer the item, then probe its reasoning at each stage. Do not merge these into one prompt. The published finding is that they catch different things — merging them throws away the reason to run both.
Step 3 — Include control items you know are clean
This is the step most teams skip, and it is the one that makes the output interpretable. Salt your batch with several items you are confident are well-designed — validated scale items from an established survey work well. If the model flags problems on your controls, its false-positive rate on this instrument is too high to trust, and you should change the prompt or the model before reading any of the real flags. The Survey Futures design used 7 European Social Survey items for precisely this purpose.
Step 4 — Vary the persona, then look at disagreement
Run the cognitive-interviewing pass with several distinct personas spanning the education, domain-familiarity, and cultural range of your real sample. You are not simulating a representative sample — the evidence is clear that LLM respondents are not psychometrically valid stand-ins for humans. You are looking for a narrower signal: items where the simulated interpretation changes across personas. Interpretive instability is a legitimate red flag regardless of whether any individual simulated answer is accurate.
Step 5 — Triage the flags and route them
Sort flags into three bins. Fix now: unambiguous wording defects the model identified with a concrete rewrite. Test with humans: anything in the high-stakes row of the triage matrix, plus every item where personas disagreed. Ignore: flags on control-comparable items where the model is pattern-matching to a taxonomy category without a plausible failure mechanism. Write the ignore decisions down; they are the audit trail that keeps the screen honest.
The probe-neutrality checklist
This is where AI pretesting carries a specific, underdiscussed methodological risk, and it deserves its own control. Conrad and colleagues found in "Probing in Cognitive Interviews can Promote Acquiescence" (2025) that when interviewers asked directly about specific unlikely interpretations of a question, respondents affirmed those interpretations more than five times as often as respondents given open-ended probes — an affirmation bias strongest among younger and less-educated respondents.
An AI moderator's probes are pre-scripted and therefore structurally prone to being directive. Worse, an LLM simulating a respondent is a compliance-tuned system being asked leading questions — the conditions for affirmation bias are doubled. Check every probe before it ships:
- Does the probe name a specific interpretation? Rewrite it as open ("What were you thinking about when you answered that?").
- Does it embed the answer you expect ("Did you include part-time work in that number?")? Replace with a paraphrase request ("Tell me in your own words what that question was asking.").
- Does it ask the respondent to evaluate the question rather than report their process? Question critique invites politeness; process reports do not.
- Are you probing every respondent identically? Consistency is the one thing an AI moderator does better than a human — it is only an advantage if the script is neutral to begin with.
- For simulated cognitive interviews specifically: run at least one probe set that offers no interpretation at all, and compare its yield to your guided set. A large gap is a warning that you are manufacturing findings.
Where Qualitati fits
Qualitati is an AI user research platform for product, UX, and customer insights teams. Two parts of it are directly relevant to pretesting.
First, conversational surveys with AI-driven follow-ups and branching logic. A conversational survey is not itself a pretest, but it changes what a pretest has to protect against: when the instrument can ask a follow-up, a merely ambiguous item is less catastrophic than in a fixed questionnaire, because the follow-up can recover the intent. Pretesting effort concentrates on the stem and the response options rather than on anticipating every misreading.
Second, AI-moderated interviews in text and voice are how you run the human half of the funnel at a cost that makes step 3 of the protocol feasible. Cognitive interviews on 8 people, in 3 languages, in a day is a different economic proposition than scheduling a moderator. Qualitati supports research in 10 languages, which matters for the cross-cultural adaptation case that Adhikari et al. studied.
What Qualitati does not do is decide for you whether an item is safe to field. The triage matrix above puts human judgment at exactly the points where the published evidence says LLM screening is weakest.
Limitations and when not to use this
Be honest about four things.
The 75% figure is a ceiling from one study, not a benchmark you inherit. It was the best of several prompt variants, on a purpose-built set of 20 questions with known planted flaws, under a specific model. The authors themselves report that detection moved with prompt design and model choice. Your hit rate on your instrument is unknown until you measure it with control items.
Planted flaws are easier than real ones. A deliberately embedded double-barrel is a cleaner target than the subtle, domain-specific misreading that emerges when a real oncologist reads a question written by a marketer. Expect worse performance on the problems you most need to find.
Simulated respondents are not respondents. The psychometric evidence is unambiguous that LLM samples are not a drop-in replacement for human survey data. Everything above treats the simulation as a detector of question defects, never as a source of substantive findings.
Do not use AI-only pretesting for clinical or health outcome instruments, legally or regulatorily consequential items, questions on sensitive experiences where the risk is participant harm rather than measurement error, or any instrument whose validity you will need to defend to a reviewer or an ethics board. In those cases the LLM pass is a first draft cleanup, and human cognitive interviewing remains the evidence.
Human-review note: the methodological claims in this article are drawn from the cited studies. Teams adopting simulated cognitive testing for regulated or high-stakes instruments should have a survey methodologist review the protocol before it replaces any part of an existing pretesting process.
Bottom line
AI survey pretesting earns its place as a screen, not a verdict. The 2026 evidence supports a specific claim — a guided LLM cognitive-testing pass finds roughly three quarters of embedded question flaws for a few dollars — and refuses a broader one, because that rate is unstable across prompts and models and the method misses different things than expert review does. Run the LLM across the whole instrument, salt it with controls so you can read the output, and spend your human interviews on the items where being wrong would cost you the study.
FAQ
Can AI replace cognitive interviewing?
No. Sturgis, Roberts and Robinson (2026) position LLM-based testing as a complement, useful for early-stage screening of large item sets and where resource constraints make human testing impossible. Simulated cognitive testing and expert review detect different problem types, and both differ from what real respondents surface.
How many questions can I pretest with an LLM?
Effectively all of them. The cost structure is the point: a full evaluation of a 20-item questionnaire completed in under an hour for a few dollars in the Survey Futures study. The binding constraint moves from budget to your capacity to triage the flags.
Which model should I use for AI survey pretesting?
Test before you commit. The published finding is that structured prompts working well on open-weight models performed poorly on the proprietary model tested, and that open-weight models detected more problems but discriminated poorly between flawed and clean items. Use control items to measure your own configuration rather than inheriting someone else's.
What is the difference between simulated expert review and simulated cognitive interviewing?
Expert review evaluates the question text against a taxonomy of known design flaws, from a methodologist's standpoint. Cognitive interviewing simulates a respondent attempting to answer and probes their reasoning. The first catches presuppositions and leading framing; the second catches difficulties that only appear when someone tries to answer.
Does AI pretesting work for multilingual surveys?
It is a promising use case, and the one with the largest resource gap. Adhikari et al. (CUI '25) tested LLM adaptation of questionnaires across US and South African samples and found modest gains in clarity and perceived bias. Treat it as a first pass; cultural referents and idiom sit in the low-detectability column of the triage matrix.
Is a conversational survey a substitute for pretesting?
No, but it changes the risk profile. Adaptive follow-ups can recover from an ambiguous item in a way a fixed questionnaire cannot, which shifts pretesting effort toward the question stem and response options. It does not fix leading framing or a wrong reference period.
Get started
Run the protocol above on your next instrument, then take the flagged items into real interviews. Start free with 30 credits — no credit card required — or view transparent pricing to see per-credit usage rates before you commit.
Last updated: September 8, 2026. This article is an independent editorial summary of publicly available research and does not represent the views of the cited authors or institutions.