研究大模型协作中欺骗与网络结构的影响,发现通信漏洞可被利用
Byzantine Cheap Talk: Adversarial Resilience and Topology Effects in LLM Coordination Games

- 通过4人猎鹿博弈测试不同模型在欺骗和拓扑限制下的行为
- 720次实验显示多数模型受骗后仍持续合作,无法集体调整策略
- 揭露通信渠道和拓扑信息暴露是潜在安全风险,适合关注AI协作安全的研究者
多智能体大模型系统依赖通信协议进行协作,但其在对抗性与结构约束下的鲁棒性尚不明确。基于先前研究表明廉价对话可促成大模型协作,我们在6种模型家族、720次试验的4人猎鹿博弈中分析两类漏洞:第一,当拜占庭代理伪装合作却背叛时,非拜占庭代理虽能在一轮内察觉背叛,却无法集体适应——大量模型持续合作,因博弈的全票收益结构而难以恢复协调;第二,显式限制通信拓扑会瓦解协作,而相同限制若隐式存在则几乎不影响协作。这表明协作失败源于代理对隐藏信息的元推理,而非信息丢失本身。我们识别出两种稳定行为模式:易叛变模型在背叛后永久转向对抗;持续合作模型则以显著个人代价维持合作。这些发现揭示了具体安全隐患:通信通道可成为对抗性注入向量,且向代理披露网络拓扑即使无攻击者也会削弱协作。
原文摘要 · Abstract (English)
Multi-agent LLM systems increasingly rely on communication protocols for coordination, yet their robustness under adversarial and structural constraints remains poorly understood. Building on prior work showing that cheap-talk channels enable cooperation in LLM coordination games, we investigate two vulnerability classes in a 4-player Stag Hunt across six model families and 720 trials. First, when Byzantine agents signal cooperation but defect, non-Byzantine agents detect the betrayal within one round yet fail to adapt collectively: a substantial fraction continue cooperating despite repeated exploitation, unable to recover coordination due to the game's unanimity payoff structure. Second, explicitly restricting communication topology collapses cooperation, while applying identical restrictions silently preserves near-perfect cooperation. This establishes that coordination failure stems from agents' meta-reasoning about hidden information, not information loss itself. We identify two stable behavioral archetypes that replicate across all model cohorts: Defection-Prone models that switch permanently after betrayal, and Cooperation-Persistent models that continue cooperating at significant individual cost. These findings reveal concrete security vulnerabilities: communication channels can be exploited as adversarial injection vectors, and disclosing network topology to agents can degrade coordination even without any adversary present.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。