Synthetic Data Accuracy: Why Digital Twin Claims Disagree
Qualitati Research Team · 2026-10-08 · 7 min read
Short answer: Published accuracy figures for digital twins and synthetic respondents range from near-perfect to near-chance mostly because studies measure different things. A 2026 paper by Netzer and Sambandam shows aggregate metrics can look strong with almost no personal data, and proposes a test that flags, before any fieldwork, which survey questions a twin can answer reliably.
If you have read vendor decks or academic papers on synthetic data, you have probably seen claims that digital twins are 95% accurate, 75% accurate, or barely better than guessing. All three numbers have published support. In "Synthetic Data in Marketing Research: How to Evaluate and When to Trust" (arXiv, September 2026), Oded Netzer (Columbia Business School) and Rajan Sambandam (TRC Insights) argue that the useful question is not whether synthetic respondents work, but when. This summary walks through their accuracy taxonomy, their empirical test on a nationally representative survey, and what it means for anyone deciding whether to put synthetic answers in front of a client.
Why do digital twin accuracy claims disagree so much?
Mostly because they use different yardsticks. According to Netzer and Sambandam (2026), a population average is far easier to predict than which individual will score higher than another, yet both get reported as "accuracy." An LLM can learn roughly where a question's average sits from its training data without knowing anything about the respondent.
The paper gives a pointed example from the Twin-2K-500 mega-study (Toubia et al. 2025): twins prompted with each person's full 500-question record scored 0.748 accuracy, twins given 14 demographic variables scored 0.746, and "twins" with no personal information at all scored 0.734. On a bounded rating scale, a constant guess near the middle is never far from anyone's answer, so a headline "75% accurate" can be reached with very little data, or none.
Normalizing by test–retest reliability has a similar blind spot. The authors note that interview-grounded agents in Park et al. (2026) reach 65.7% raw accuracy, reported as 82.6% once divided by participants' own consistency, while demographics-only agents reach 74% on that same normalized scale.
What are the four families of accuracy measures?
The paper sorts the literature's metrics into four families, ordered from least to most demanding of the data. The key distinction is whether a measure needs matched individual records and whether it scores the level of an answer or only the ordering of people.
| Family | What it compares | Needs matched individuals? | Main risk |
| 1. Cross-question correspondence | Question-level means correlated across many items | No | Near-perfect scores from population priors alone |
| 2. Question-level distributions | Synthetic vs human means, spreads, choice shares, WTP | No | Misses that twins are under-dispersed |
| 3. Matched individual accuracy | Each twin's answer vs its own person's answer | Yes | High information-free floor on bounded scales |
| 4a. Within-person, across questions | Correlation across items for each person | Yes | Rewards good averages with no individuation |
| 4b. Within-question, across people | Whether twins rank people correctly on one question | Yes | Ignores level shifts and compressed spread |
Which accuracy level does a research decision need?
It depends on the decision tier. Netzer and Sambandam map common commercial uses to three tiers:
- Population-level reads (concept screening, average willingness-to-pay, demand curves): Family 2 validity against a human benchmark, plus a check on dispersion, not only means.
- Heterogeneity (segmentation, targeting, differentiated pricing): at minimum Family 4b sorting plus distributional checks within segments.
- Individual-level uses (personalization, twin-as-panelist, respondent-level imputation): Family 3 and/or 4b.
Their practical advice to buyers: ask which tier a vendor's accuracy claim was established at, and treat a Family 1 statistic as validation only for a Family 1 question.
What is the "forgotten question" problem?
It is the question a team realizes it should have asked only after fieldwork closed. Re-fielding is slow and costly, so even an imperfect low-cost estimate has value. The authors argue this is where twins are most promising: the researcher already holds a full human study on the same population and category, and the twin only needs to extend each respondent's record by one item.
How did the study test digital twins?
The authors used a 2025 survey of 3,063 respondents from a nationally representative panel, run for Rice University's C-CUBES center and fielded by TRC Insights. It covered 108 attitude questions (six value drivers in each of 18 sectors, rated 1–10) plus 20 demographic and background questions. In a leave-one-out design, GPT-5.4-mini answered each held-out question for each respondent, conditioned either on demographics only or on demographics plus the other five attitude items from the same sector.
The results, according to Netzer and Sambandam (2026):
- Demographics-only twins reached an individual-level correlation (Family 4b) of just r = 0.08 with real answers.
- Adding same-category attitudes raised it to r = 0.57 and cut individual error from 2.01 to 1.46 scale points.
- Yet the aggregate error was lower for the demographics-only twins (0.39 vs 0.48), a clean illustration of how a topline metric can hide the absence of individual signal.
- Even with attitudes included, 25.9% of questions had a twin–human correlation below 0.5.
Can you tell in advance which questions a twin can answer?
Yes, at least partly. The paper's answerability diagnostic fits a random forest that predicts the twins' answers from the exact data placed in their prompts. Its out-of-bag R² shows how much of the twin's output is driven by the respondent's record rather than the model's priors, and it needs no human answers to the forgotten question.
Average R² was 0.14 for demographics-only twins and 0.65 when attitudes were added. Within the richer condition, R² correlated r = 0.61 with realized twin–human correlation. Screening at R² above 0.7 kept 46 of 108 questions (43%), raised mean correlation to 0.65, cut individual error to 1.19, and reduced poorly answered questions from 25.9% to 4.3%. The authors note the cutoff should be set on a hold-out sample in practice.
Two cheaper screens also carried signal. Embedding similarity between the forgotten question and fielded items correlated r = 0.53 with twin accuracy, though its gains were smaller. Sixteen TRC researchers rating 18 questions matched realized accuracy at r = 0.77 and the R² screen at r = 0.87, which the authors flag as promising but imprecise given the small question set.
What this means for researchers
The paper's checklist translates directly into practice for anyone evaluating synthetic respondents or a digital twin panel:
- State the decision first, then pick the accuracy family it requires; report every family your data supports.
- Compare every number to an information-free floor (empty persona, demographics only) and, where possible, a test–retest ceiling.
- Run a "wrong-person twin" test: build each twin from someone else's record and measure the gap.
- Validate on novel items, since well-known scales may sit in the training data.
- Report accuracy across respondent groups; twins tend to fit educated, higher-income and moderate respondents better.
- Do not attach conventional standard errors to synthetic estimates; query counts are a design choice, not a sample.
The broader lesson: accuracy belongs to a twin–question pair, not to the twin. Keeping humans first, collecting rich same-category data in your surveys, and using twins only for short, screened extrapolations is the pattern this evidence supports.
FAQ
Are digital twins accurate enough to replace survey respondents?
Not as a general substitute. In this study, twins grounded in same-category data reached a mean individual-level correlation of 0.57 with real answers, and about a quarter of questions fell below 0.5 without screening.
Why can a twin with no personal data still score 73% accuracy?
On bounded rating scales, predictions near the population average are never far from any one answer. The paper recommends comparing results to empty-persona and demographics-only baselines, and using balanced accuracy for multiple-choice items.
What is the R² answerability screen?
A random forest predicts each twin's answer from the data in its prompt. A high out-of-bag R² means the twin is using the respondent's record; a low one means it is falling back on general knowledge. It is a necessary, not sufficient, condition for accuracy.
Where can I read the original paper?
The preprint is Netzer and Sambandam (2026), "Synthetic Data in Marketing Research: How to Evaluate and When to Trust", arXiv:2609.13995.
Last updated: October 8, 2026
This article is an independent editorial summary of third-party research by the Qualitati Research Team. It is not affiliated with or endorsed by the paper's authors. Figures are reported as published in the arXiv preprint; consult the original for full methods and results.