Can LLMs Be Digital Twins of Real People? A 2026 Study
Qualitati Research Team · 2026-07-06 · 8 min read
Short answer: Not yet — but closer than most researchers assume. A June 2026 study by Kinzinger and Hartmann built LLM "individual-level twins" from real German survey microdata and predicted held-out answers with up to 78.8% best-case accuracy, while rank-order correlation with real respondents topped out around r = 0.59. That is useful for exploration and pre-testing, not a replacement for real people.
What did the study test?
The study asked whether an LLM, conditioned on a real person's background data, can answer survey questions the way that specific individual actually would. According to Kinzinger and Hartmann (2026), the authors constructed individual-level twins from the German Socio-Economic Panel (SOEP), a long-running representative household survey, and then measured how well each twin reproduced that person's answers to questions the model had never seen.
This is a sharper test than most synthetic-respondent work. Rather than asking whether an LLM can imitate a demographic group in aggregate, it asks whether the model can mimic a named individual's specific response — the harder, more honest bar for anyone hoping to use synthetic panels in place of real recruitment.
How the twins were built and evaluated
The authors ran a systematic evaluation rather than a single prompt, which is what makes the numbers worth citing. According to Kinzinger and Hartmann (2026), the design crossed several factors:
- 3 open-weights language models — so results are not tied to one proprietary system.
- 5 levels of "information depth" — how much of each person's background the twin was given, from sparse to full.
- 2 embedding approaches — comparing raw dialog history against condensed narrative summaries.
- 2 reasoning modes — with and without explicit "thinking" before answering.
Crossed together, that produced more than 2.1 million evaluated responses across 500 individuals and 183 held-out questions. The held-out design matters: the twin is judged on questions it was not conditioned on, which guards against the model simply parroting information it was already handed.
How accurate were the digital twins?
The best configurations were surprisingly strong on point accuracy but weaker on preserving individual differences. The headline figures from Kinzinger and Hartmann (2026):
| Metric |
Result |
What it tells you |
| Best-cell accuracy |
78.8% |
Under the best factor combination, twins matched the real held-out answer most of the time. |
| Rank-order correlation (Fisher-z) |
r ≈ 0.590 |
Moderate agreement on how people differ from each other — the harder, more valuable signal. |
| Best information source |
Raw dialog history |
Feeding the model raw response history beat feeding it tidy narrative summaries. |
| Effect of explicit reasoning |
Better correlation, flat accuracy |
"Thinking" mode improved rank-order fidelity without moving point accuracy. |
Read together, these say the twins are good at guessing the modal answer but only moderately good at capturing what makes one respondent different from the next. For research that lives on distributions and subgroup differences, that gap is the whole story.
What did the study conclude?
The authors argue that the constraint has shifted. According to Kinzinger and Hartmann (2026), "twin-based market research is no longer gated by data design, but by item volume, model selection," and construction choices — that is, the limiting factors are now how many questions you ask, which model you pick, and how you feed in a person's history, rather than whether individual-level twins are feasible at all. Two practical findings stand out: raw dialog histories beat narrative summaries, and explicit reasoning improves rank-order correlation without changing raw accuracy.
What this means for researchers
Individual-level LLM twins are a legitimate tool for pre-testing and exploration — not a substitute for human respondents in confirmatory work. Three implications:
- Use twins to pilot, not to conclude. An accuracy near 79% and a correlation near 0.59 are strong enough to sanity-check a questionnaire, surface obvious framing problems, or rank candidate concepts — but not to certify an effect size you would defend in a report.
- Grounding beats invention. Twins built on real microdata outperform free-floating personas. This echoes prior work on why ungrounded synthetic personas drift toward a bland average (see our note on synthetic user persona collapse).
- Guard the tails. Because twins capture the modal answer better than individual variation, they will under-represent minority views and edge cases — exactly the segments qualitative research exists to hear.
If you want to experiment with grounded synthetic respondents alongside real ones, Qualitati's Digital Twin Panel and Synthetic Focus Group are built for exactly this pre-testing role — generate first-pass reactions from persona-grounded twins, then validate the ones that matter with real AI-moderated interviews.
Frequently asked questions
Can an LLM accurately predict how a specific person will answer a survey?
Partially. In the 2026 Kinzinger and Hartmann study, LLM twins grounded in real survey microdata reached 78.8% best-case accuracy on held-out questions but only moderate rank-order correlation (r ≈ 0.59), meaning they capture the typical answer better than what makes each individual distinct.
Are synthetic respondents good enough to replace real survey panels?
No. The evidence supports using synthetic or twin-based respondents for exploration, questionnaire pre-testing, and concept ranking — not for final, decision-grade estimates, where they under-represent minority and edge-case views.
Why did raw dialog history beat narrative summaries?
Summaries compress away the specific, idiosyncratic detail that distinguishes one respondent from another. Feeding the model a person's raw response history preserved that signal, which improved fidelity across model configurations at full information depth.
What is an individual-level digital twin?
It is an LLM conditioned on one real person's background and prior responses so it can stand in for that specific individual, as opposed to a persona that only represents a demographic group in the aggregate.
Last updated: July 6, 2026. This article is an independent editorial summary of third-party research; figures and claims are drawn from the cited preprint (Kinzinger & Hartmann, 2026) and have not been independently verified by Qualitati.