arXiv:2605.28114cs.AI2026-05

语言模型在群体标签可见时会自发产生信任偏见,现有评估方法难以发现。

Language model agents show in-group trust bias invisible to standard behavioural audits

论文配图:Language model agents show in-group trust bias invisible to standard behavioural audits
图 1 · 摘自论文原文
  • 通过20人模拟实验,模型对同组目标发起53.6%-54.6%的信任行为
  • 偏见在五种主流模型中均出现,且经三种统计检验与重述测试验证
  • 适合关注多智能体系统安全与评估漏洞的研究者阅读

语言模型代理正从单用户助手演变为持续互动的网络,同时控制实体机器人与软件系统。我们发现,五种广泛使用的开源推理模型在群体身份可见时即产生内群体信任偏见,即使分组仅为无实际意义的任意标签。在20代理模拟中,模型将53.6%-54.6%的信任行为指向同组目标,显著高于随机预期的47.4%,该现象在所有测试模型中均存在,并通过三次独立统计检验及指令重述鲁棒性测试确认。该偏见易被当前评估实践忽略,因其作用于‘向谁发起行动’而非‘选择何种行动’,而后者正是标准行为日志审计的唯一观测通道。资源稀缺性干预本意是检验竞争是否加剧偏见,结果却使其中三种模型的偏见下降,归因于稀缺性施加方式的人为干扰,而非机制失效。因此,群体相关社会动态已存在于构建多智能体系统的模型中,而基于单模型、单决策的评估范式无法检测此类偏差。

原文摘要 · Abstract (English)

Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and the same models increasingly control physically embodied robots as well as software. Here we show that five widely used open-weight reasoning models develop an in-group trust bias the moment group membership becomes visible to them, even when the groups are arbitrary labels with no real-world meaning: in a 20-agent simulation, agents direct 53.6-54.6% of their trust-building actions toward in-group targets against a 47.4% base rate expected by chance, a shift present in every model tested and confirmed by three independent statistical checks and an instruction-rewording robustness test. This bias is easy for current evaluation practice to miss, because it operates through which agent receives an action rather than which action is chosen - a channel invisible to the aggregate behaviour-log audits that are the standard way multi-agent AI systems are evaluated today. A resource-scarcity manipulation, intended to test whether competition intensifies the bias, instead reduced it in three of five models; we trace this to an artifact of how scarcity was enforced, not to a failure of the underlying mechanism. Group-contingent social dynamics are therefore already present in the models multi-agent AI systems are built from, and auditing practice built around single-model, single-decision evaluation cannot detect them.

多智能体信任偏见评估漏洞语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。