用SILICA测试大模型代理社会是否真像人类,结果多数不符。
Benchmarking large language model agent societies against human behavioural distributions

- 设计SILICA工具,三重检验模型行为可靠性
- 仅8/11模型首轮合作达标,终态无一匹配人类
- 模型依赖记忆而非真实互动,适合研究者验证
大型语言模型代理群体正被广泛用作实验社会。但存在三大疑虑:代理行为是否类人、结果是否对装置微调敏感、观察到的社会动态是否源于真实互动而非模型记忆。本文提出SILICA——一个开放测试工具,涵盖五个含已发表人类基准数据的环境,每个环境配以规则不变但结构扰动的变体,以及收益方向偏离记忆结果的版本。十二个开源权重模型在单张消费级显卡上运行。结果显示,仅有8/11模型在首轮公共品贡献上落在等效区间内,终态贡献与人类协作区间均无匹配。仅改变两个动作的呈现顺序,即导致一个模型合作分下降58点。固定报价序列测试表明,仅唯一经推理训练的模型将接受阈值置于激励要求位置;另两个部分调整,两个反向移动,三个从未形成阈值。惯例形成依赖对名称的共享先验,而非协商,但一旦先验被打破,协商重现。根据本文定义的认证阶梯,当前硅基社会仅支持探索性结论,不支持因果推断。
原文摘要 · Abstract (English)
Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。