arXiv:2605.08268cs.MAcs.AI2026-05被引 1

研究语言模型协作中内鬼如何拖延共识,提出用学习模型+强化学习实现高效攻击。

Insider Attacks in Multi-Agent LLM Consensus Systems

论文配图:Insider Attacks in Multi-Agent LLM Consensus Systems
图 1 · 摘自论文原文
  • 用世界模型学习良性代理行为状态,构建攻击代理的仿真环境。
  • 训练出的攻击者使良性共识率下降,分歧持续时间更长。
  • 适合关注多智能体安全与对抗性语言交互的研究者阅读。

大型语言模型(LLMs)正被广泛部署于多智能体系统中,智能体通过自然语言交流协同完成任务。此类系统的关键能力是共识形成,即智能体迭代交换信息并更新决策以达成一致结果。然而,现有大多数多智能体LLM框架假设所有参与智能体均与系统目标对齐。实际上,恶意内鬼可能作为合法成员参与,却暗中追求对抗性目标。本文研究多智能体LLM共识系统中的内鬼操纵问题,将其形式化为一个序列决策任务:恶意智能体试图延迟或阻止良性智能体达成一致。为使攻击优化可计算,我们提出基于世界模型的框架,学习良性智能体潜在行为状态的替代动态,并利用该模型训练攻击者。初步结果显示,训练后的攻击者比直接恶意提示基线更能降低良性共识率并延长分歧时间。这些结果表明,将隐式世界模型与强化学习结合,是语言驱动多智能体系统中自适应内鬼攻击的有前景方向。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in multi-agent systems where agents communicate in natural language to solve tasks jointly. A key capability in such systems is consensus formation, where agents iteratively exchange messages and update decisions to reach a shared outcome. However, most existing multi-agent LLM frameworks assume that all participating agents are aligned with the system objective. In practice, a malicious insider may participate as a legitimate member of the group while pursuing a hidden adversarial goal. In this work, we study insider manipulation in multi-agent LLM consensus systems. We formalize the problem as a sequential decision-making task in which a malicious agent seeks to delay or prevent agreement among benign agents. To make attack optimization tractable, we propose a world-model-based framework that learns surrogate dynamics over the latent behavioral states of benign agents and then trains an attacker using reinforcement learning based on this learned model. Preliminary results show that the trained attacker reduces the benign consensus rate and prolongs disagreement more effectively than the direct malicious-prompt baseline. These results suggest that combining latent world models with reinforcement learning is a promising direction for adaptive insider attacks in language-based multi-agent systems.

多智能体安全攻防共识系统语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。