Synthetic Participants and Digital Twins in Research: Uses, Limits, and Ethics (2026)
Written by Fengming Liu (UCL) · Reviewed by Prof. Shubin Yu (HEC Paris) · May 2026 · Updated: 2026-05-30 · 16 min read
Imagine pre-testing your interview guide on a panel of participants who never tire, respond instantly, and cost nothing — before you recruit a single real person. That is the appeal of synthetic participants. It's also where researchers most need a clear head, because the same technology that makes a useful rehearsal tool can quietly manufacture findings that look real but aren't.
This guide explains synthetic participants and digital twins for researchers: what they are, how they're built, the legitimate uses, the hard limits, and the ethics. It builds on our GenAI-Assisted Qualitative Research guide, which covers the broader workflow.
What are synthetic participants and digital twins?
Synthetic participants are AI-generated personas — powered by large language models — that simulate how a person with a given background might respond to research questions. Digital twins are a specific, data-grounded kind: AI personas built from a real individual's own interview transcripts or documents, so the twin answers new questions in a way calibrated to that person. Both let researchers query a "participant" without recruiting or re-recruiting humans.
The crucial distinction is their grounding. A generic synthetic persona is constructed from a demographic description and the model's training data; a digital twin is anchored in real data from a specific person. Twins tend to be more faithful — but neither is a substitute for fresh human data.
Key concepts
- Synthetic participant / AI persona — a model prompted to role-play a described person or segment.
- Digital twin — a persona derived from a real individual's prior responses, used for follow-up questions.
- Synthetic focus group — multiple AI personas discussing a topic together to surface diverse perspectives and tensions before fieldwork.
- "Silicon sampling" — the contested idea of using LLM-generated responses as a stand-in for survey or interview samples.
How are they created?
There are two broad routes:
- From a persona specification — you describe demographics, attitudes, and context, and the model role-plays that person. Fast and flexible, but only as grounded as the description and the model's priors.
- From real data (digital twins) — you supply a participant's actual transcripts or documents, and the system builds a twin that answers new questions consistently with what that person actually said. This is the more defensible approach for follow-up research, and platforms increasingly attach a calibration score that estimates how closely the twin matches the real source.
Where they genuinely help
- Piloting instruments — pressure-test an interview guide or survey for confusing or leading questions before fieldwork.
- Scoping hypotheses — explore the rough shape of a question space to sharpen your design.
- Rehearsing moderation — practice probing and flow against a synthetic focus group.
- Idea generation — surface angles, tensions, and counter-arguments you hadn't considered.
- Follow-up on existing data (twins) — ask new questions of digital twins built from completed interviews without re-contacting participants, then validate the most promising findings with real follow-ups.
Where they fall short — and the risks
| Risk | What goes wrong |
| Not real evidence | Synthetic responses reflect a model's priors, not lived experience; they can't discover the genuinely unexpected. |
| Bias and homogenization | Models flatten diversity and can stereotype groups, especially under-represented ones. |
| Manufactured consensus | Personas may agree too readily, inventing a false sense of agreement. |
| Hidden hallucination | Fluent, confident answers can be entirely fabricated. |
| Ethical misrepresentation | Passing synthetic data off as human data is a research-integrity violation. |
Synthetic participants are a rehearsal space, not a stage. They can help you prepare, pilot, and explore — but conclusions about real people must rest on data from real people.
Digital twins vs. generic synthetic personas
| Aspect | Digital twin | Generic synthetic persona |
| Grounding | A real person's own data | A description + model priors |
| Fidelity | Higher, and measurable via calibration | Lower and harder to verify |
| Best use | Follow-up questions on existing participants | Early ideation and piloting |
| Main risk | Over-trusting the twin beyond its source | Stereotyping and hallucination |
Validity and calibration
The central question is always: how well does the synthetic response match what a real person would say? Calibration addresses this by comparing twin or persona outputs against held-out real data and scoring the agreement. A high calibration score builds confidence for exploratory use; it never licenses treating synthetic output as primary evidence. Always validate consequential findings against real participants.
Ethical use and disclosure
- Never misrepresent — synthetic data must be labeled as synthetic in any report or publication.
- Consent for twins — building a digital twin from someone's data should be covered by informed consent.
- Guard against bias — check that personas don't caricature the groups they represent.
- Disclose methods — state how and where synthetic participants were used, including models and prompts.
- Keep humans central — use synthetic methods to support real research, not to avoid engaging real people.
Best practices
- Use synthetic participants for preparation and exploration, not as final evidence.
- Prefer data-grounded digital twins over generic personas for anything consequential.
- Calibrate and validate against real data before trusting a result.
- Label and disclose all synthetic data transparently.
- Watch for bias and false consensus in every output.
Tools
Some AI research platforms now build these methods in directly. QualiTaTi's Synthetic Focus Group runs multi-agent persona discussions to explore a question before fieldwork, and its Digital Twin Panel extracts personas from completed interviews or documents and attaches calibration scoring — so follow-up questions stay anchored to real source data.
A worked example
- Design — before recruiting, run a synthetic focus group of five personas to pilot your guide and spot leading questions.
- Fieldwork — conduct 20 real interviews and analyze them normally.
- Follow-up — build digital twins from those 20 transcripts and ask a new question that emerged during analysis.
- Validate — check calibration scores, then confirm the most important twin-derived insights with a few real participants.
- Report — present real findings as primary, clearly labeling any synthetic exploration as such.
Key takeaways
- Synthetic participants are AI personas; digital twins are personas grounded in a real person's own data.
- They're valuable for piloting, scoping, rehearsal, ideation, and grounded follow-up — not as a replacement for human data.
- Risks include bias, homogenization, false consensus, and hallucination.
- Calibration estimates fidelity; consequential findings must still be validated against real participants.
- Always label synthetic data and disclose how it was used.
Frequently asked questions
What is a synthetic participant?
A synthetic participant is an AI-generated persona, powered by a large language model, that simulates how a person with a given background might answer research questions — useful for piloting and exploration, but not a source of real evidence.
What is a digital twin in research?
A digital twin is an AI persona built from a real individual's own interview transcripts or documents, allowing researchers to ask that "participant" new follow-up questions without re-recruiting them. Twins are more faithful than generic personas and can be calibrated against the source data.
Can synthetic participants replace real participants?
No. They reflect a model's priors rather than lived experience and can introduce bias, false consensus, and hallucination. They support preparation and exploration, but conclusions about real people require data from real people.
Are synthetic focus groups useful?
Yes, for rehearsal and ideation. Running multiple AI personas in discussion can surface tensions, counter-arguments, and angles before fieldwork, helping you refine your design — provided you don't treat the output as real consumer or participant data.
Is it ethical to use synthetic participants?
It can be, when used transparently: label synthetic data as synthetic, obtain consent before building digital twins from real data, check for bias, and never present synthetic responses as human findings.
What is calibration and why does it matter?
Calibration compares a synthetic persona or twin's responses against real data and scores how closely they match. It indicates how much confidence to place in synthetic output for exploratory use — but it never makes synthetic data a substitute for real evidence.