arXiv:2509.18052cs.CLcs.CY2025-09被引 18

提出6大原则,揭示大模型模拟社会行为的漏洞

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies

  • 提出PIMMUR六原则,系统评估大模型社会模拟的可靠性
  • 89.7%的研究违反至少一条原则,61.0%提示过度控制结果
  • 实验证明许多所谓'涌现'行为实为方法缺陷所致

大语言模型正被用于模拟人类集体行为,但其方法学严谨性仍缺乏探讨。通过对39项近期研究的系统审计,我们识别出六大普遍缺陷——代理设定、互动、记忆、控制、无知与现实性(统称PIMMUR)。分析显示,89.7%的研究至少违反一项原则,损害模拟有效性。我们发现,前沿大模型仅在50.8%的情况下能正确识别社会实验本质,而61.0%的提示存在过度控制,预先决定结果。通过复现五个代表性实验(如电话游戏),我们证实:当严格执行PIMMUR原则时,多数报告的集体现象消失或反转,表明许多所谓‘涌现’行为实为方法论伪象,而非真实社会动态。研究警示当前大模型模拟可能反映的是模型特异性偏差,而非普适的人类社会行为,对将大模型作为人类社会科学代理工具构成严峻挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed to simulate human collective behaviors, yet the methodological rigor of these "AI societies" remains under-explored. Through a systematic audit of 39 recent studies, we identify six pervasive flaws-spanning agent profiles, interaction, memory, control, unawareness, and realism (PIMMUR). Our analysis reveals that 89.7% of studies violate at least one principle, undermining simulation validity. We demonstrate that frontier LLMs correctly identify the underlying social experiment in 50.8% of cases, while 61.0% of prompts exert excessive control that pre-determines outcomes. By reproducing five representative experiments (e.g., telephone game), we show that reported collective phenomena often vanish or reverse when PIMMUR principles are enforced, suggesting that many "emergent" behaviors are methodological artifacts rather than genuine social dynamics. Our findings suggest that current AI simulations may capture model-specific biases rather than universal human social behaviors, raising critical concerns about the use of LLMs as scientific proxies for human society.

大模型社会模拟有效性方法论缺陷

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。