Magicgametime a family ledger
Technology

Challenges Persist for AI Surrogates in Behavioral Research

Published Sep 03, 2026 Reads 963 By Sujata Gupta

A recent study highlights the limitations of AI digital twins in accurately mimicking human behavior for social science research, indicating a need for human input.

Challenges Persist for AI Surrogates in Behavioral Research

Limitations of AI in Behavioral Studies

As researchers explore the potential of using AI to replicate human behavior in behavioral studies, a new investigation reveals significant limitations. The research, published on September 2 in Science Advances, finds that AI digital twins, designed to emulate individual behaviors, often present distorted views and create a misleading “funhouse mirror” effect. This raises immediate concerns regarding the reliability of AI systems in accurately modeling complex human behaviors.

The Experiment: Expectations vs. Reality

Olivier Toubia, a computational social scientist at Columbia Business School, expressed optimism about AI's capabilities but ultimately found the results of recent experiments disappointing. The study collected data from over 2,000 participants across the United States, who answered more than 500 questions assessing various characteristics—including demographics, personality traits, and cognitive abilities. This large-scale dataset, now downloaded around 25,000 times, was intended to propel forward research in digital twin technology, encouraging scientists to interrogate the nuances of human behavior through AI.

The expectation was that this extensive input would allow researchers to create AI representations, or digital twins, of participants. By inputting their information into a large language model (LLM), the idea was to generate simulated responses based on individual profiles. However, despite an impressive data collection effort, the results highlighted a deeper challenge in utilizing AI to mirror human cognition effectively.

Assessment of Digital Twins

Across 19 experiments, the performance of these digital twins was assessed on various topics, such as political donations and views on algorithmic hiring practices. While the twins performed better than random guesses, they were incorrect about 25% of the time, only matching the accuracy of chatbots that relied solely on demographic data. This outcome suggests that simply having more data isn't a silver bullet for achieving deeper understanding or predictive accuracy in behavioral simulations.

Interestingly, the digital twins demonstrated a greater understanding of individual differences compared to LLMs with limited information. For instance, when asked to rate traits like self-control, the AI might distinguish between a rating of 2 and 4. (And this is the part most people overlook.) Yet, despite such distinctions, the AI still missed the mark, indicating that a greater understanding of variance does not equate to a meaningful grasp of human behavior. This highlights a qualitative gap that may prove difficult to overcome through data alone.

The Underlying Issues

However, the underperformance of the digital twins stemmed from oversimplified responses that conformed to demographic stereotypes. An analysis revealed a tendency for these AI surrogates to share homogeneous views, especially among more affluent participants. This homogeneity is problematic, as it perpetuates a narrow understanding of diverse human experiences. Additionally, the twin AIs displayed biases, such as demonstrating a higher level of trust in others and insensitivity to technological risks, which starkly contrasts with the varied reactions of human respondents.

Hadi Hosseini, a Penn State researcher specializing in AI, corroborated these findings with studies on AI decision-making in health care. He noted similar distortions toward more rational decision-making, highlighting the disconnect between AI interpretations and actual human behavior. The trend suggests a critical flaw in how AI systems interpret the complexities of human emotions and choices.

Future Enhancements and Opportunities

Despite these challenges, opportunities for improvement exist. Toubia suggested that incorporating dynamic interactions, like continuous engagement with individuals, could enhance the development of digital twins. Shifting from static questions to a more fluid approach may provide richer datasets and enhance accuracy in behavioral simulations. It's not merely about asking the right questions; it’s about how those questions are presented and how responses are contextualized. If you're working in this space, the shift towards a more adaptive and responsive model could become your key to unlocking better results.

Potential Applications in Social Science Research

In exploring future applications, Toubia noted that AI could still play a role in social science research. For example, researchers might benefit from the capabilities of digital twins to generate detailed responses when human respondents are fatigued. This could reduce bias in specific scenarios where participant engagement might wane. Furthermore, pre-testing experimental designs using AI twins could streamline the research process by optimizing participant engagement and improving the quality of data collected.

Implications for Future Research

However, Toubia urges caution, emphasizing the complexity of human behavior that may elude even the most sophisticated AI. Social scientists often rely on surveys and scales to capture the essence of human experience, using qualitative and quantitative methods to peel back layers of thought and emotion. The challenge of translating the intricacies of human thought into synthetic data remains formidable. Researchers must keep realistic expectations while probing the possibilities of AI applications in social sciences.

What this means for you—if you're looking at integrating AI into your research—might entail grappling not only with the data but also with the limitations inherent within AI technology. The balance between leveraging AI for efficiency and maintaining the integrity of human experience is delicate, and findings like these remind us of the pitfalls that await in this ambitious intersection of technology and human behavior.

Source: Sujata Gupta · www.sciencenews.org

Discussion

Sign in to join the discussion.