Across 19 social-science experiments, AI “digital twins” answered as their human counterparts would only imperfectly, missing the mark about a quarter of the time and falling short of replacing live subjects.
More than 2,000 U.S. participants had provided 500-plus survey answers and test results to build the twins, yet models using that rich profile performed only about as well as chatbots given demographic data alone.
The study found the twins compressed human variation into more homogeneous, stereotype-skewed responses, while also appearing more trusting, less worried about technology and generally more rational than the people they mimicked.
Accuracy was better for wealthier and more educated participants, suggesting uneven reliability, though the richer profiles did somewhat preserve differences between individuals better than demographic-only models.
Researchers said the systems may still help with tasks like pretesting experiments or generating fuller draft responses, but argued synthetic subjects remain far from capturing real human behavior.