arXiv:2508.09230cs.MAcs.AI2025-08ICML被引 1

为多智能体系统设计免疫机制,防攻击扩散。

Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems

  • 通过分发特制治愈样本实现事前免疫与事后恢复。
  • 实验证明可显著降低感染传播率,理论保证系统鲁棒性。
  • 适合关注AI系统安全、多智能体协作的开发者与研究者。

基于视觉语言模型(VLM)的智能体是具备状态感知与环境交互能力的自主实体。多智能体系统由多个专业智能体协作完成复杂任务。其核心安全属性是鲁棒性,即在对抗攻击下仍能维持系统完整性。然而现有系统设计缺乏鲁棒性考量,单一智能体被攻破后,攻击可迅速蔓延至其他智能体,导致整个系统失效。为此,我们提出一种名为Cowpox的新防御机制,可形式化提升多智能体系统的鲁棒性。该机制采用分布式设计,通过限制攻击向其他智能体传播的期望次数,提高智能体恢复率。核心思想是在暴露前生成并分发特殊治愈样本,使智能体获得免疫能力,并帮助已感染智能体恢复。我们在实验中验证了Cowpox的有效性,并提供了理论上的鲁棒性保障。

原文摘要 · Abstract (English)

Vision Language Model (VLM)-based agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent systems comprise specialized agents who collaborate to solve a (complex) task. A core security property is robustness, stating that the system should maintain its integrity under adversarial attacks. However, the design of existing multi-agent systems lacks the robustness consideration, as a successful exploit against one agent can spread and infect other agents to undermine the entire system's assurance. To address this, we propose a new defense approach, Cowpox, to provably enhance the robustness of multi-agent systems. It incorporates a distributed mechanism, which improves the recovery rate of agents by limiting the expected number of infections to other agents. The core idea is to generate and distribute a special cure sample that immunizes an agent against the attack before exposure and helps recover the already infected agents. We demonstrate the effectiveness of Cowpox empirically and provide theoretical robustness guarantees.

多智能体系统安全免疫机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。