AI digital twins — large language models configured to answer questions as one specific, real, named human being — have quietly become one of the most attractive ideas in social science. Feed a model everything a person has ever told you about themselves, the pitch goes, and it will answer the next question the way they would. A new benchmark published in Science Advances on 2 September 2026 tested that pitch across nineteen pre-registered studies and found it does not hold. The authors, led by Tianyi Peng at Columbia, describe what they built as funhouse mirrors that systematically distort human behavior.
That phrase is doing precise work. A funhouse mirror is not broken. It reflects a real person, recognisably, in the right general shape — and it stretches, flattens and smooths them in ways that are consistent enough to look like a rule and wrong enough to ruin any measurement you take from the reflection. That is what the study found across 164 separate outcomes, from hiring decisions to misinformation sharing to green energy defaults.
This article walks through the benchmark itself, the five distortions the researchers named, the numbers behind each one, the competing evidence from a well-known Stanford project that reached a friendlier conclusion, and what any of it means if someone is currently trying to sell you synthetic survey respondents. If you follow the research that changes how AI actually gets used in practice, we track it in our AI models and tools hub as it lands. The short version: the method is real, the promise is early, and the error is not random noise you can average away.
Table of contents
- What an AI Digital Twin Actually Is
- The Benchmark That Broke the AI Digital Twin Story
- Distortion One: The AI Digital Twin Does Not Individuate
- Distortion Two: Stereotyping Inside the AI Digital Twin
- Distortion Three: Representation Bias in AI Digital Twin Accuracy
- Distortion Four: The Ideological Tilt an AI Digital Twin Cannot Shake
- Distortion Five: Hyper-Rationality in Every AI Digital Twin
- The Distribution Problem Behind Every AI Digital Twin
- What Sixty-Eight Experts Got Wrong About AI Digital Twin Performance
- Does a Better Model Fix the AI Digital Twin?
- Where the AI Digital Twin Evidence Genuinely Disagrees
- How to Use an AI Digital Twin Without Being Fooled
- Frequently Asked Questions About the AI Digital Twin Study
- Conclusion: A Mirror That Flatters and Flattens
- References and Further Reading
What an AI Digital Twin Actually Is
The term is borrowed, and the borrowing causes confusion. It is worth separating the two meanings before looking at any of the results.
The industrial twin versus the behavioural one
A digital twin in engineering is a live simulation of a physical asset — a turbine, a production line, a building — fed by sensor telemetry and used to predict wear or test changes safely. That discipline is mature, standards-backed and largely a data-plumbing problem, which is why we treated it as one in our guide to digital twin development.
An AI digital twin of a person is a different animal entirely. There is no telemetry, no ground-truth physics and no sensor feed. There is only a transcript of what someone once said about themselves, loaded into a language model’s context window as a persona, and a hope that the model will extrapolate correctly.
How the twins in this study were built
The researchers did not use a thin persona. Each AI digital twin was conditioned on that individual’s own prior answers to more than 500 survey questions — roughly 128,000 characters of self-report per person, covering personality inventories, attitudes, demographics and behavioural history.
That is an unusually rich brief. If a language model can imitate a specific person at all, this is close to the best-case setup: no guessing from a name and a postcode, but a dossier the person wrote themselves.
Why researchers want this to work
The appeal is straightforward and not cynical. Human subjects are expensive, slow to recruit, prone to fatigue, and in some study designs exposed to real psychological distress. A synthetic panel is instant, free at the margin, infinitely re-runnable and raises no ethics-board objection.
Marketing teams want the same thing for the same reasons. Test forty pack designs against a simulated customer base overnight, or run a pricing study without fielding one. That is the commercial pull behind every AI digital twin pitch currently circulating, and it is why a rigorous negative result matters.
The Benchmark That Broke the AI Digital Twin Story
Most prior work on simulated respondents ran one study, on one topic, with one outcome. The Columbia team’s contribution is scale and pre-registration.
Nineteen studies, 164 outcomes, 1,784 people
The team ran nineteen pre-registered sub-studies spanning 164 distinct outcomes, comparing 1,784 real participants against their corresponding twins, with more than 13,000 human observations across the substudies. The outcomes were deliberately varied: attitudes toward hiring algorithms, intentions to share misinformation, privacy preferences, default effects in green energy enrolment, evaluations of creative ideas and perceptions of algorithmic fairness.
Breadth is the point. A single study that happens to work proves nothing about a method; 164 outcomes across nineteen designs produce a distribution you can actually reason about.
The correlation that decides everything
Across all of it, the average correlation between an AI digital twin and the human it was built from was r = 0.20. That is a weak relationship. It is not zero — the twins are not random — but it is nowhere near the level at which you would substitute a simulated answer for a real one in a study you intended to publish or a decision you intended to fund.
The paper’s own summary is that twins are “only modestly more accurate than those of a homogeneous base LLM.” That comparison is the uncomfortable part, and it deserves its own section.
The baselines that should worry you
The researchers ran three controls. A twin with the full 500-question persona scored r = 0.197. A twin given nothing but the person’s demographics scored r = 0.145. A twin given an empty persona — no individual information at all — scored r = 0.080. Random responses scored r = 0.001.
At the individual level the picture is starker still. Full-persona accuracy came in at 0.748 against a demographics-only figure of 0.746 — a difference that did not reach significance (p = 0.37). Half a thousand personal answers bought roughly nothing over knowing someone’s age, income and education.
Distortion One: The AI Digital Twin Does Not Individuate
The first of the five named distortions is the one that explains most of the rest. Twins are not different enough from one another.
Under-dispersion in 154 of 164 outcomes
For any given question, the spread of answers across a real population is wide. The spread across their twins is narrower. The researchers found the standard deviation of twin responses was lower than the human standard deviation in 154 of 164 cases — 93.9%, and the variance reduction was statistically significant in 140 of those.
An AI digital twin, in other words, regresses toward a consensus answer. It produces the reasonable middle of a population rather than the population.
Twins resemble each other more than they resemble people
The sharper version of this finding uses mean absolute deviation. The gap between a full-persona twin and an empty-persona twin averaged 0.175. The gap between a full-persona twin and the actual human it was modelling averaged 0.252, significantly larger (p < 0.01).
Read that again slowly. Two twins built from two completely different people’s 500-answer dossiers are more similar to each other than either is to the person it was supposed to represent. The persona moves the model less than the person differs from the model’s default.
Why sameness is worse than noise
Random error averages out. If your simulated respondents were wrong in unpredictable directions, a large enough synthetic sample would still recover the population mean. Systematic compression does not behave that way — it shrinks the tails, hides minorities of opinion and makes every segment look more agreeable than it is.
That failure mode is particularly bad for the commercial use cases. Concept testing exists to find the people who hate the idea. An AI digital twin panel that smooths dissent will tell you every idea is fine.
Distortion Two: Stereotyping Inside the AI Digital Twin
If the persona is not driving the answers, something else is. The second distortion identifies what.
Demographics do almost all the work
The mean absolute deviation between a full-persona twin and a demographics-only twin was just 0.132 — substantially smaller than the 0.252 gap between the full-persona twin and the human. The paper’s conclusion is blunt: answers from full-persona twins are closer to those from demographics-only twins than to those from either empty-persona twins or real humans.
The model, handed a rich individual dossier, appears to compress it back down into a demographic category and answer from that.
What that means in practice
This is a stereotyping mechanism, and it is worth being precise about why it is a problem rather than merely a limitation. A demographic prior is not useless — it carries real signal. But a system marketed as modelling this person while actually modelling people like this person misrepresents what it is doing.
It also imports whatever demographic associations sit in the training corpus, which is a well-documented source of bias in any large language model application. The persona provides a veneer of individuation over a group-level guess.
The 500-question dossier bought very little
The clearest single number in the study is the individual-level accuracy comparison: 0.748 with the full dossier, 0.746 with demographics alone, p = 0.37. Everything the participants disclosed about themselves beyond their basic categories moved the needle by two-thousandths of a point.
| Persona given to the model | Correlation (r) | Individual accuracy |
|---|---|---|
| Full persona, 500+ prior answers | 0.197 | 0.748 |
| Demographics only | 0.145 | 0.746 |
| Empty persona | 0.080 | 0.734 |
| Random responses | 0.001 | 0.629 |
Distortion Three: Representation Bias in AI Digital Twin Accuracy
The third distortion concerns who the twins get right, and it is the one with the clearest equity implications.
Accuracy tracks education and income
The researchers found systematic differences in twin accuracy across demographic groups. Twins were more accurate for participants with higher education and higher income, and less accurate for people outside moderate political and religious-attendance patterns.
The direction is not surprising once stated. A language model’s training corpus over-represents the writing of educated, affluent, online people, so its default persona sits closer to them.
The compounding effect on synthetic samples
Consider what happens if a research team replaces a nationally representative panel with an AI digital twin panel. The simulated sample is not uniformly degraded. It is accurate for the groups already best served by research and least accurate for the groups usually under-served by it.
That converts a general accuracy problem into a distributional one. Findings drawn from such a panel would be most wrong precisely where policy attention tends to be most needed.
Why the raw average hides it
An aggregate correlation of 0.20 says nothing about who sits above and below it. This is the standard argument for disaggregated evaluation, and it is the same discipline the NIST AI Risk Management Framework asks for in any consequential AI deployment: measure performance by subgroup, not just in total.
Distortion Four: The Ideological Tilt an AI Digital Twin Cannot Shake
The fourth distortion is the most quotable, because the twins lean in two directions at once.
Pro-human and pro-technology simultaneously
The study found twins were consistently more trusting of people than the real individuals they modelled were, and simultaneously more accepting of technology — including, pointedly, more accepting of algorithmic hiring than the humans they were imitating.
Those two tilts do not obviously belong together, which is part of what makes them look like an artefact of the model rather than a property of the people.
The self-favouring judgement
One result is worth isolating. When asked to evaluate creative ideas, twins rated human-generated ideas lower than ideas that had come from twins. A simulated respondent that prefers simulated output is not a neutral instrument for any study about attitudes to AI.
The twins also under-reported platform usage relative to their humans, another consistent directional gap rather than scatter.
Why this contaminates AI research specifically
The single most tempting use of an AI digital twin panel is to survey attitudes toward AI itself — it is fast-moving, commercially urgent and expensive to field. It is also the exact topic where a documented pro-technology tilt makes the instrument unusable.
That problem generalises. Any research question where the model holds a stable opinion is a question the twins will answer with that opinion, wearing the participant’s name.
Distortion Five: Hyper-Rationality in Every AI Digital Twin
The fifth distortion is the one an outside researcher summarised most sharply.
Perfect knowledge where humans guess
On factual questions, twins displayed near-perfect knowledge, deviating from the human participants who — being human — got things wrong. They also selected normative, higher-cognitive-ability answers more often than the people they represented.
Hadi Hosseini of Penn State, who was not involved in the work, put the direction plainly to Science News: “LLMs distort human judgment in a very specific direction toward something that is very rational [and] more reasonable than actually what people are.”
The bounded-rationality effects vanish
The most technically damning finding is that the twins failed to replicate the attraction effect and the compromise effect — two of the best-established demonstrations that human choice is context-dependent and not strictly rational.
Those effects are not obscure. They are load-bearing results in behavioural economics and marketing. A synthetic panel that cannot reproduce them cannot be used to study choice architecture, which is one of the fields most eager to adopt the method.
Fatigue, effort and essay-length answers
There is a mundane version of the same problem. Tired humans give one-sentence answers late in a long survey. An AI digital twin never tires and will happily produce an essay for question 400. That difference alone introduces a systematic gap in response style before anything substantive is measured.
| Distortion | What it does | Supporting figure |
|---|---|---|
| Insufficient individuation | Twins cluster near a consensus answer | Under-dispersed in 154 of 164 outcomes |
| Stereotyping | Answers driven by demographic category | MAD 0.132 vs demographics-only twin |
| Representation bias | Better accuracy for higher education and income | Systematic subgroup accuracy gaps |
| Ideological bias | More trusting of people and of technology | Rated human ideas below twin ideas |
| Hyper-rationality | Too knowledgeable, too normative | Attraction and compromise effects absent |
The Distribution Problem Behind Every AI Digital Twin
Individual accuracy is one question. Whether a synthetic panel reproduces a population is another, and the study is harsher on the second.
105 of 164 outcomes were significantly different
At the distribution level, twin responses differed significantly from human responses on 105 of 164 outcomes — 64%. The average gap was 0.352 standard deviations.
A third of a standard deviation is not a rounding error. On a typical seven-point attitude scale it is the difference between a finding and a null result.
Why aggregation does not rescue it
The usual defence of synthetic respondents is that individual predictions do not need to be good if the aggregate is. This benchmark removes that defence. The aggregate is also wrong, in a consistent direction, on nearly two-thirds of the outcomes tested.
Anyone building an AI digital twin panel for market sizing or predictive analytics should treat that as the headline number rather than the correlation.
Half the experimental effects
The paper also notes the twins replicated only about half the experimental effects observed in humans, with a suspicious asymmetry: effects from previously published paradigms came through more strongly than effects from novel ones. That pattern points at training-data leakage rather than genuine behavioural fidelity.
What Sixty-Eight Experts Got Wrong About AI Digital Twin Performance
One of the study’s neatest components had nothing to do with the models.
The forecasting exercise
The researchers asked 68 human experts to predict how the twins would perform. The experts did poorly. In three of five test cases, the twins’ treatment effects fell outside the experts’ 95% confidence intervals.
Domain expertise did not confer the ability to guess which questions a language model would answer like a person. That is a strong argument for benchmarking every intended use rather than reasoning about it.
Intuition is not a substitute for measurement
This is the practical takeaway for anyone considering the method. You cannot look at a proposed study and judge whether an AI digital twin will handle it. The failures are not where informed people expect them.
The authors’ response was to release their full dataset and code as a standardised testbed, so that future claims can be checked rather than asserted.
The testbed and the dataset behind it
The underlying resource is Twin-2K-500, a dataset of more than 2,000 people answering more than 500 questions each, published in Marketing Science and hosted openly on Hugging Face. Olivier Toubia of Columbia Business School, a co-author, told Science News it had been downloaded around 25,000 times.
Does a Better Model Fix the AI Digital Twin?
The obvious objection is that the study used the wrong model, and that the next generation will close the gap. The researchers tested that directly.
What they tried
The default configuration was GPT-4.1 at temperature 0.7. Alternatives included GPT-5, DeepSeek, Gemini 3 Pro, a version of GPT-4.1 fine-tuned on the Twin-2K-500 data, and Centaur, a model fine-tuned specifically for cognitive prediction.
None of them changed the conclusion. The best configuration found anywhere in the sweep was GPT-4.1 at temperature 0, which reached a correlation of 0.232 — an improvement on 0.197, and still weak.
Why scale is not the missing ingredient
The distortions the paper documents are not obviously capability failures. Under-dispersion, demographic compression and hyper-rationality all look like consequences of how models are aligned and decoded, not gaps in what they know.
A model trained to be helpful, harmless and reasonable is being asked to be idiosyncratic, occasionally uninformed and sometimes irrational. Those objectives are in tension, and more parameters do not resolve tension.
The honest framing from the authors
Toubia’s summary to Science News was measured rather than dismissive: “There’s some promise. But [the twins’ performance] was overall a bit disappointing.” His conclusion for buyers was firmer — “We need to be realistic in terms of the expectations we have from synthetic data.”
Where the AI Digital Twin Evidence Genuinely Disagrees
There is a well-known result pointing the other way, and any fair reading has to account for it.
Stanford’s interview-grounded agents
A Stanford-led team including Joon Sung Park and Michael Bernstein built agents from two-hour semi-structured interviews with 1,052 Americans. On held-out General Social Survey items, their interview-grounded agents reached 83% of participants’ own two-week test-retest consistency, rising to 86% when interviews and surveys were combined, against 74% for demographics-only agents.
That sounds far better than r = 0.20. The two studies are measuring different things.
The benchmark is the difference
The Stanford figure is expressed as a fraction of how well people replicate their own answers a fortnight later — a deliberately forgiving yardstick that absorbs human inconsistency. The Columbia benchmark reports raw correlation against the human’s actual response, and adds an explicit comparison against a persona-free baseline.
Note too that the demographics-only baseline in the Stanford work sat at 74% against 83% for interviews — a real gap, but the same qualitative story: category information carries most of the weight.
What both studies agree on
Both find that demographics-only agents are surprisingly competitive. Both find accuracy disparities across racial and ideological groups. And both point at the same open question: whether richer, more dynamic input can pull an AI digital twin away from the model’s default persona.
| Study | Columbia, Science Advances 2026 | Stanford, arXiv 2024-26 |
|---|---|---|
| Participants | 1,784 compared to twins | 1,052 Americans |
| Persona source | 500+ survey answers | Two-hour interview transcript |
| Scope | 19 studies, 164 outcomes | GSS, Big Five, economic games |
| Headline metric | Raw correlation, r = 0.20 | 83–86% of test-retest consistency |
| Demographics baseline | r = 0.145, near-identical accuracy | 74% of test-retest consistency |
| Verdict | Not ready for primetime | General-purpose simulation is feasible |
How to Use an AI Digital Twin Without Being Fooled
The paper is not an argument for abandoning the method. It is an argument for describing it accurately.
Advisor, not carbon copy
The authors offer their own reframing: “Perhaps use cases that consider digital twins as well-informed advisors are more promising than use cases that consider digital twins as carbon copies.” That is a real distinction with practical consequences.
An advisor is useful for generating hypotheses, drafting question wording, spotting an obvious flaw in a stimulus or exploring a space before you spend money. A carbon copy is what you need before you replace a panel, and that is the claim the evidence does not support.
Questions worth asking a vendor
If a research supplier is offering synthetic respondents, four questions separate a defensible offer from a marketing one. What is your correlation against held-out human responses, not against another model? What is your accuracy for the least-represented subgroup in the sample? Does your synthetic panel reproduce the variance of the real one, or only its mean? And which published behavioural effects does it replicate?
Any vendor who cannot answer the variance question has not looked at the failure mode this study spent nineteen experiments documenting.
Where the method is defensible today
Pre-testing survey instruments, generating candidate hypotheses, stress-testing question wording and pilot work ahead of real fielding are all reasonable. So is using twins as one signal inside a broader data science programme rather than as the evidence base itself.
What is not defensible is presenting synthetic results as population estimates, or using them for any question where the model has a stable opinion of its own — which, for AI attitudes, it demonstrably does. The same care applies to the broader problem of telling machine output from human output, which we covered when Pangram’s chief executive argued that detection is harder than “real or fake”.
Frequently Asked Questions About the AI Digital Twin Study
What is an AI digital twin of a person?
It is a large language model prompted with a specific individual’s own prior answers, personality data and demographics, then asked new questions as though it were that person. It is unrelated to the engineering sense of the term, which simulates a physical asset from sensor data.
How accurate were the twins in this study?
The average correlation with the humans they modelled was r = 0.20 across 164 outcomes, and individual-level accuracy was 0.748 against 0.746 for a twin given only demographics. At the distribution level, twins differed significantly from humans on 64% of outcomes.
Why is low variance a bigger problem than low accuracy?
Random error cancels out across a large sample; systematic compression does not. Twin answers were under-dispersed in 94% of outcomes, which means a synthetic panel understates disagreement, hides minority views and makes every segment look more moderate than it is.
Would a newer model fix this?
The team tested GPT-5, Gemini 3 Pro, DeepSeek, a fine-tuned GPT-4.1 and Centaur. The best result anywhere was r = 0.232. The distortions look like consequences of alignment and decoding rather than missing capability, so more scale is unlikely to be the answer on its own.
Does this mean synthetic respondents are useless?
No. The authors describe twins as promising but “not yet ready for primetime” and suggest treating them as well-informed advisors rather than carbon copies. Pilot work, instrument testing and hypothesis generation remain reasonable uses.
Where can I check the study myself?
The paper appeared in Science Advances on 2 September 2026 under DOI 10.1126/sciadv.aeh8260, with a preprint on arXiv as 2509.19088. The underlying Twin-2K-500 dataset and the evaluation code are public.
Conclusion: A Mirror That Flatters and Flattens
The funhouse-mirror framing is the most useful thing to take from this work, because it names the specific failure rather than a vague one. These systems are not producing noise. They are producing a consistent, directional deformation: everyone a little more similar, a little more rational, a little more trusting, a little more comfortable with technology, and a little more like the demographic average they were sorted into.
That is precisely the kind of error that survives averaging and looks like a finding. An AI digital twin panel will hand you clean, plausible, well-written results with the disagreement quietly removed — and nineteen pre-registered studies now say those results are significantly different from the humans on nearly two-thirds of the questions asked.
The honest position is the one the authors took: real promise, an open benchmark, and a clear instruction not to deploy it as a replacement yet. The mirror works. It just does not tell you what you look like.
References and Further Reading
Digital Twins are Funhouse Mirrors: Five Systematic Distortions (arXiv preprint)
Science Advances: Digital twins are funhouse mirrors (DOI 10.1126/sciadv.aeh8260)
Science News: AI agents not ready to replace humans in behavioral research
Tech Xplore: AI digital twins struggle to predict human behavior
Twin-2K-500: A dataset for building digital twins of over 2,000 people
LLM-Digital-Twin datasets on Hugging Face
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
NORC: The General Social Survey
NIST AI Risk Management Framework
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.