LLM能模拟人类行为做心理实验,但效果夸大且敏感话题易出错。
Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management
- 用三款顶级LLM复现156项心理学实验,测试其模拟人类反应能力。
- 主效应复制率73%-81%,交互效应46%-63%,但效应值高出人类2-3倍。
- 在种族、性别等敏感话题上表现差,零结果常被误判为显著,需谨慎使用。
人工智能正越来越多地融入社会科学科研,尤其在理解人类行为方面。大型语言模型(LLMs)在模拟心理实验中的人类反应方面展现出潜力。本研究大规模复现了来自顶级社会科学期刊的156项心理实验,采用三种先进LLM(GPT-4、Claude 3.5 Sonnet、DeepSeek v3)。结果显示,尽管LLMs在主效应复制上表现良好(73%-81%),交互效应中等至较强(46%-63%),但其效应值普遍高于人类研究,Fisher Z值约为人类研究的2-3倍。值得注意的是,涉及种族、性别与伦理等社会敏感话题的研究,其复制率显著降低。当原始研究报告零结果时,LLMs产生显著结果的比例高达68%-83%——虽可能因数据更干净、置信区间更窄,但仍提示效应值高估风险。研究证明,LLMs在心理研究中兼具潜力与挑战,可作为快速假设验证和预研工具,但复杂社会现象与文化敏感问题仍需人类参与与审慎解读。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) is increasingly being integrated into scientific research, particularly in the social sciences, where understanding human behavior is critical. Large Language Models (LLMs) have shown promise in replicating human-like responses in various psychological experiments. We conducted a large-scale study replicating 156 psychological experiments from top social science journals using three state-of-the-art LLMs (GPT-4, Claude 3.5 Sonnet, and DeepSeek v3). Our results reveal that while LLMs demonstrate high replication rates for main effects (73-81%) and moderate to strong success with interaction effects (46-63%), They consistently produce larger effect sizes than human studies, with Fisher Z values approximately 2-3 times higher than human studies. Notably, LLMs show significantly lower replication rates for studies involving socially sensitive topics such as race, gender and ethics. When original studies reported null findings, LLMs produced significant results at remarkably high rates (68-83%) - while this could reflect cleaner data with less noise, as evidenced by narrower confidence intervals, it also suggests potential risks of effect size overestimation. Our results demonstrate both the promise and challenges of LLMs in psychological research, offering efficient tools for pilot testing and rapid hypothesis validation while enriching rather than replacing traditional human subject studies, yet requiring more nuanced interpretation and human validation for complex social phenomena and culturally sensitive research questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。