Why AI Survey Surrogates Overstate What People Like (2026)
Qualitati Research Team · 2026-08-19 · 6 min read
LLM-generated survey respondents say they like almost twice as many things as real people do, and the relationships between their preferences barely resemble human ones. A 2026 study of 277,470 synthetic responses found silicon samples reproduce coarse population averages while losing nearly all of the structure researchers actually analyze.
What did the study test?
The study asked whether LLM "surrogates" can stand in for human survey respondents on cultural taste. Ma, Zhang, Ang and Chen (2026), in Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates, took the U.S. Survey of Public Participation in the Arts (SPPA) — a long-running national survey covering 1982–2022 — and used its 9,249 human respondents as ground truth.
For each real respondent, the authors generated 30 synthetic answers conditioned on that person's demographic profile, producing 277,470 silicon responses across three models: GPT-4.0, Claude 3.7 Sonnet, and DeepSeek R1. Each surrogate was asked which of 17 music genres it liked. Because the human baseline is a probability sample with known demographic structure, the comparison can be made at three separate levels: how much people like, what goes with what, and who likes what.
How accurate were the synthetic respondents?
Accuracy depended entirely on which level you look at. At the crudest level — the overall share of genres liked — silicon samples were in the right ballpark but systematically too enthusiastic. According to Ma et al. (2026), real SPPA respondents reported liking 24.18% of the 17 genres, while the synthetic samples reported liking 45.89% (OpenAI), 46.14% (DeepSeek), and 39.77% (Anthropic).
That is roughly double the human rate, and the gap widened for less popular genres: for non-popular genres, the OpenAI surrogates ran +32.29 percentage points above the human figure of 26.52%. The authors describe this as a systematic positive bias for liking, producing inflated estimates of cultural omnivorousness — and note it cannot be explained away as ordinary WEIRD (Western, educated, industrialized, rich, democratic) sampling bias.
Three levels of failure
| Level of analysis | Human benchmark | Silicon result | Verdict |
| Marginal rates (how much people like) | 24.18% of genres liked | 39.77%–46.14% | Directionally plausible, badly inflated |
| Relationality (what goes with what) | Bootstrap benchmark near r = 1.00 | Pearson's r = 0.06 | Essentially no correspondence |
| Demographic association (who likes what) | Observed SPPA effects | Inflated and sign-flipped effects | Misleading for subgroup analysis |
Why does the "relationality" result matter most?
The relationality finding is the one that should stop a synthetic-panel project. Taste rarely matters in isolation; researchers care about which preferences travel together, because that co-occurrence structure is what segments, personas, and typologies are built from.
The correlation between silicon and human taste associations was 0.06, against a bootstrap benchmark close to 1.00 — meaning a genuinely random split of the human sample reproduces the structure almost perfectly, while the LLM version does not reproduce it at all. Individual pairings were reshuffled rather than merely noisy: the blues–classic rock association dropped from 0.40 to 0.04, while electronic–gospel rose from 0.06 to 0.57. The models did not fail to find structure; they invented a different one.
Do the demographic patterns survive?
No — and they fail in both directions, which is worse than uniform error. The authors report age effects biased negative for 14 of the 17 genres, a race effect on reggae of −0.365 representing roughly a nine-fold inflation of the true association, and a gender effect on Broadway of +0.453 against a human value of +0.107.
The practical consequence: any analysis that crosses a synthetic panel by age, gender, class, or race is reading amplified or reversed effects. A synthetic panel might tell you a segment exists, at a magnitude that would clear conventional significance thresholds, when the human data says otherwise.
What this means for researchers
Treat synthetic respondents as a tool for generating hypotheses and stress-testing instruments, not for producing estimates you will report or act on. Four working rules follow from this study:
- Never take marginal rates at face value. A doubled liking rate would distort any market-sizing or incidence estimate built on it.
- Do not run segmentation on synthetic data. Co-occurrence structure was the study's weakest level, and segmentation is built directly on it.
- Do not report subgroup differences from synthetic panels. Demographic effects were inflated and sometimes sign-flipped.
- Validate against a real sample before you rely on anything. The study's design — real survey as ground truth, synthetic responses generated per-profile — is a template any team can reuse at small scale.
The highest-value uses that survive this evidence are the pre-fieldwork ones: piloting question wording, checking that a screener behaves, rehearsing an interview guide before it meets real participants. That is how we frame our own Digital Twin Panel — a rehearsal surface, not a replacement for fieldwork. When you need the real structure of what people believe and why, it still has to come from people, which is why the recommended path is running an AI Interviewer study with human participants and analyzing what they actually said.
Note also the scope condition: this was cultural taste, a domain where preferences are socially patterned and status-laden. Whether LLM surrogates do better on factual or behavioral items is an open question this study does not answer — but the burden of proof now sits with anyone claiming they do.
FAQ
Can AI-generated survey respondents replace human samples?
Not for estimation. The 2026 SPPA study found synthetic respondents roughly doubled the rate of liking (24.18% human vs. 39.77%–46.14% synthetic) and reproduced almost none of the correlation structure between preferences (r = 0.06).
Which models were tested?
GPT-4.0, Claude 3.7 Sonnet, and DeepSeek R1. All three showed the positive-liking bias, so this is not a single-vendor artifact.
What is silicon sampling?
Silicon sampling is prompting an LLM with a demographic profile and asking it to answer survey questions as that person would, producing synthetic respondents. This study generated 30 surrogates per real respondent.
Is there any legitimate research use for synthetic respondents?
Yes — pre-testing questionnaires, checking survey logic and screeners, and generating hypotheses to test with real participants. The failure documented here is in using synthetic output as evidence about a population.
Source: Ma, X., Zhang, M., Ang, S., & Chen, M. (2026). Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates. arXiv:2606.30085.
Last updated: 19 August 2026. This is an independent editorial summary of third-party research; QualiTaTi is not affiliated with the authors.