为AI代理实验建立预注册机制,提升研究可信度
Preregistration for Experiments with AI Agents

- 提出针对AI代理实验的预注册模板
- 识别出模型选择、提示词等八大自由度漏洞
- 适合关注实验可复现性的研究者参考
大型语言模型和自主AI代理的兴起催生了“仿真”行为实验的新方法论范式。这类实验最初用于以AI代理模拟人类在认知、决策与社会动态中的行为,如今因AI代理越来越多地代表个人与组织进行协商与决策,理解其行为已成为独立的研究重点。尽管此类实验在可扩展性、成本效率和实验控制方面具有前所未有的优势,但也继承甚至放大了传统人类被试研究中长期存在的方法学缺陷。本文主张,应将人类实验中关键的预注册实践扩展至AI代理实验。系统梳理了实验中引入的研究者自由度,如模型选择、提示词设计、参数设置及结果导向的重设计等,并指出低迭代成本与缺乏报告规范使这些选择易于被滥用且难以察觉。为此,提出专为AI代理实验设计的预注册模板,呼吁会议、期刊与资助机构将预注册作为该新兴研究范式的标准要求。
原文摘要 · Abstract (English)
The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: "in silico" behavioral experiments. Originally conceived as a way to use AI agents as proxies for human participants in studies of cognition, decision-making, and social dynamics, this approach has taken on new significance -- as AI agents increasingly negotiate, transact, and make consequential decisions on behalf of people and organizations, understanding their behavior has become a research priority in its own right. While these experiments with AI agents offer unprecedented advantages in terms of scalability, cost efficiency, and experimental control, they also inherit, and in some cases amplify, methodological vulnerabilities that have long plagued human subjects research. To address these issues, this paper argues that preregistration practices -- central to improving the credibility of human subjects experiments -- should now be extended to experiments with AI agents. We systematically catalog the researcher degrees of freedom that experiments with AI agents introduce -- model selection, prompt wording, settings, and outcome-contingent redesign, for example -- and show how the low cost of iteration and lack of reporting norms make these choices both easy to exploit and difficult to detect. We propose a preregistration template tailored to experiments with AI agents and call on conferences, journals, and funding agencies to make preregistration standard practice for this emerging research paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。