Can Richer Personas Make LLMs Human Surrogates? A 2026 Test
Qualitati Research Team · 2026-10-01 · 7 min read
Short answer: No. A 2026 working paper by Ahn, Mao and Lee tested LLM human surrogates on four datasets covering more than 400,000 participants and found that models reproduce each survey item's average answer well, but explain only 3.05% of how individual respondents deviate from that average. Richer personas and fine-tuning did not close the gap.
Synthetic respondents are usually sold on one premise: give the model enough about a person, such as demographics, personality scores or a long interview, and it will answer the way that person would. A new working paper, Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates (Daehwan Ahn, University of Georgia; Chengfeng Mao, MIT Sloan; Dokyun Lee, Boston University; arXiv, August 2026), tests that premise directly. Its answer is that most of the apparent accuracy of LLM surrogates comes from knowing which questions people tend to answer high or low, not from knowing anything about the individual.
What did the study test?
The study tested whether LLM surrogates can predict individual answers once the easy part, each item's average human response, is taken out. According to Ahn, Mao and Lee (2026), the analysis used four datasets so the result would not hinge on one survey:
| Dataset | Respondents | Items / outcomes | Role in the study |
| Megastudy (Peng et al. digital-twin data, 18 studies) | 1,784 | 160 | Primary test |
| SocSci210 (210 studies) | >400K | 5,998 | Robustness |
| Twin-2K-500 Survey | 2,058 | 126 | Robustness + test-retest benchmark |
| ANES (American National Election Studies) | 4,270 | 11 | Robustness (original "silicon samples" setting) |
The key move is simple. For every question, the authors subtracted the human average from both the human answer and the LLM answer. What remains is each respondent's personal deviation: whether this person sits above or below the crowd on this specific question. A true individual surrogate should predict that deviation. A model that only knows the item averages will not.
How accurate are LLM survey surrogates at the individual level?
They are barely accurate at all once item averages are removed. In the primary Megastudy analysis, LLM item means correlated strongly with human item means (Pearson r = 0.78 across 133 items). But the pooled de-meaned R², the share of respondent-specific variation the LLM explains, was 3.05%. The ceiling set by human test-retest reliability, meaning how consistently the same people answer the same items twice, was 53.6%. The LLM reached only 5.7% of what is achievable.
The authors also compared the LLM against a deliberately dumb baseline: predict each person's answer as the average of everyone else's answers to that item. That baseline uses no information about the person. Across 1,631 respondents, the LLM's mean per-person correlation was 0.34; the leave-one-out human average scored 0.45. In other words, the persona-prompted model did worse than simply reporting what other people said.
Do richer personas or fine-tuning help?
No tested variant closed the gap. The paper reports these de-meaned R² values (Ahn, Mao and Lee, 2026, Table 2):
- Full-persona prompt (Megastudy): 3.05%
- Strongest non-fine-tuned prompt (Megastudy): 4.44%
- Fine-tuned GPT-4.1 (Megastudy): 2.31%, below the best plain prompt
- Socrates-Qwen2.5-14B, fine-tuned on SocSci210: 7.57% on studies it had seen, 0.73% on held-out studies
- Twin-2K-500 Survey: 3.87% to 6.21% across specifications
- ANES four-item set: 10.77%, the highest result, concentrated in party identification and ideology
Personas are not useless. When the authors shuffled personas across respondents 10,000 times, the shuffled versions never exceeded 0.028%, so correct matching does carry some real signal. It is simply far too little to treat the output as a stand-in for that person.
Why don't personas fix the problem?
Because the missing information is not who a person is in general, but how they react to a particular question. Using generalizability theory, the authors split prediction error into parts. Stable person effects, the kind of thing a persona summary describes, accounted for only 4.9% of error variance. The person-by-item interaction, how a specific respondent departs from the average on a specific item, accounted for 44.0%, about 8.9 times larger. A persona is fixed across all questions, so it can encode the first kind of signal but not the second.
Do synthetic responses look like real response distributions?
Not closely. LLM answers were consistently narrower than human ones. Median item-level standard deviations were 65% of the human value in the Megastudy, 50% in SocSci210 and 57% in the Survey. Models also used fewer response categories, with median ratios between 43% and 73% of the human figure. The shape of the distribution differed too: shape accounted for 31% to 40% of the distance between human and LLM distributions, so the mismatch is more than a shifted mean or a squeezed spread. The authors call the overall pattern item-mean surrogacy.
How should researchers validate a synthetic respondent claim?
Match the test to the claim. The paper proposes four tests that a surrogate must pass before it can be described as standing in for individuals:
- Recover deviations, not means. Report de-meaned accuracy, not just correlation with item averages.
- Generalize to new items and tasks. Accuracy on seen studies does not count; the Socrates result shows how fast it falls on held-out ones.
- Beat shuffled personas. If swapping personas between respondents does not hurt prediction, the model is producing a generic answer, not simulating a person.
- Reproduce the distribution. Check spread, number of categories used and shape, not only the mean.
The authors also report a literature check that explains why this matters. Of 63 papers since January 2023 that propose, deploy or test LLMs as human surrogates, 53 were advocates. Among the 44 advocates that offered empirical validation, 33 validated only at the aggregate level and just 4 reported individual-level metrics.
What this means for researchers
Synthetic respondents remain useful for aggregate, early-stage work: spotting items likely to hit a ceiling, catching gross design problems in a questionnaire and forming rough expectations before fieldwork. The paper does not support using them to replace individual respondents, segment-level claims built on individual simulation, or any analysis that depends on spread and disagreement.
In practice that suggests a workflow. Use a digital twin panel or a synthetic focus group to stress-test a discussion guide or survey draft, then run the real study with people and compare. If you report synthetic results, state the fidelity level you tested, and run at least the shuffle and distribution checks above. Our earlier summary of persona prompting research reaches a compatible conclusion from a different dataset.
Limitations of the study
The authors are explicit about scope. The analysis covers ordered, bounded survey responses, so open-ended answers and behavioral tasks may behave differently. It does not test systems that retrieve a respondent's full answer history at prediction time, or models trained specifically for individual prediction. The measures are linear, so nonlinear individual structure could be missed. And it is a working paper on arXiv, not yet peer-reviewed; the preregistration is due to be released on publication.
FAQ
What is item-mean surrogacy?
It is the pattern in which an LLM surrogate reproduces the average human answer to each question but fails to predict how individual respondents differ from that average, and also produces narrower, differently shaped response distributions.
Are LLM digital twins accurate?
At the aggregate level they can look accurate, with item-mean correlations around 0.78 in this study. At the individual level, after removing item averages, they explained 3.05% of respondent-specific variation against a 53.6% human reliability ceiling.
Does fine-tuning make synthetic respondents more realistic?
Not in this study. A fine-tuned GPT-4.1 scored below the best plain prompt, and a model fine-tuned on SocSci210 dropped from 7.57% on seen studies to 0.73% on held-out ones.
Can I still use synthetic respondents?
Yes, for piloting and aggregate expectations. Treat them as a draft of what the average respondent might say, then validate with real participants before drawing individual or segment-level conclusions.
Last updated: October 1, 2026
This article is an independent editorial summary of third-party research by the Qualitati Research Team. Qualitati is not affiliated with the authors. For full methods and results, read the original paper.