arXiv:2602.00685cs.AI2026-02被引 4

用可复现的实验框架测试大模型作为模拟参与者的表现

HumanStudy-Bench: Towards AI Agent Design for Participant Simulation

  • 将模拟参与者设计为完整实验流程的智能体,结合基础模型与行为设定
  • 在12项经典研究中验证,6000+次试验结果与真人高度一致
  • 提供可自由探索的设计空间和双指标评估体系,适合社会科学研究者

大型语言模型越来越多地被用作社会科学实验中的模拟参与者,但其行为常不稳定且对设计选择敏感。以往评估常混淆模型基础能力与实验实例化效果,难以区分结果是源于模型本身还是代理配置。本文将‘参与者模拟’视为全实验协议上的代理设计问题,其中代理由基础模型与行为假设规范(如参与者属性)共同定义。我们提出HUMANSTUDY-BENCH,一个面向实践者的开放平台与执行引擎,支持开发与评估针对特定实验场景的代理。该平台通过人机协同的筛选-提取-执行-评估流程,端到端重建已发表的人类实验,保留原始刺激、条件与统计方法,同时允许自由探索代理设计空间。我们引入两项互补指标,分别量化代理在统计显著性结论与效应量上与人类的一致性,并考虑人类参考数据中的有限样本不确定性。与社会科学家合作,我们在涵盖个体认知、策略互动与社会心理学的12项基础研究中进行了验证,覆盖6000+次试验。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base model capabilities with experimental instantiation, obscuring whether outcomes reflect the model itself or the agent setup. We instead frame $\textbf{participant simulation as an agent-design problem}$ over full experimental protocols, where an agent is defined by a base model and a specification (e.g., participant attributes) that encodes behavioral assumptions. We introduce $\textit{HUMANSTUDY-BENCH}$, an open platform and execution engine designed for practitioners to develop and evaluate agents tailored to their target experimental settings. The platform reconstructs published human-subject experiments via a human-in-the-loop Filter--Extract--Execute--Evaluate pipeline that preserves the original stimuli, conditions, and statistical procedures end to end, while allowing practitioners to freely explore the agent design space. We introduce two complementary metrics that quantify agreement with humans on both the significance conclusion and the effect size, while accounting for finite-sample uncertainty in the human reference data. In collaboration with social scientists, we validate the platform on a suite of 12 foundational studies covering 6,000+ trials across individual cognition, strategic interaction, and social psychology

AI代理实验模拟社会实验语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。