arXiv:2603.01213cs.MAcs.LG2026-03被引 6

测试大模型代理在无利益冲突下的共识能力,发现协作易失败。

Can AI Agents Agree?

  • 用全对称模拟测试语言模型代理的共识行为
  • 群体越大、越有恶意节点,达成一致越困难
  • 主要问题在于进程卡死而非结果被篡改,适合关注协作可靠性者阅读

大型语言模型正越来越多地作为协作代理使用,但其在对抗性共识场景中的表现尚未系统研究。本文通过同步全对称模拟,在标量值的拜占庭共识游戏中评估基于LLM的代理。实验设定无利益偏好,仅关注是否达成一致。在数百次模拟中,涵盖不同模型规模、群体大小和拜占庭比例,结果表明:即使在良性环境下,有效共识也不可靠,且随群体增大而恶化;引入少量恶意代理后成功率进一步下降。失败主要源于活跃性丧失,如超时和收敛停滞,而非细微的数值篡改。整体表明,当前LLM代理群体在无利益设置下仍不具备可靠的共识能力,警示依赖稳健协作的部署应用。

原文摘要 · Abstract (English)

Large language models are increasingly deployed as cooperating agents, yet their behavior in adversarial consensus settings has not been systematically studied. We evaluate LLM-based agents on a Byzantine consensus game over scalar values using a synchronous all-to-all simulation. We test consensus in a no-stake setting where agents have no preferences over the final value, so evaluation focuses on agreement rather than value optimality. Across hundreds of simulations spanning model sizes, group sizes, and Byzantine fractions, we find that valid agreement is not reliable even in benign settings and degrades as group size grows. Introducing a small number of Byzantine agents further reduces success. Failures are dominated by loss of liveness, such as timeouts and stalled convergence, rather than subtle value corruption. Overall, the results suggest that reliable agreement is not yet a dependable emergent capability of current LLM-agent groups even in no-stake settings, raising caution for deployments that rely on robust coordination.

多智能体共识机制大模型协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。