arXiv:2605.13213cs.AI2026-05中稿 · CVPR被引 1

提出分层攻击框架,揭示多模态多智能体系统的安全弱点。

Hierarchical Attacks for Multi-Modal Multi-Agent Reasoning

论文配图:Hierarchical Attacks for Multi-Modal Multi-Agent Reasoning
图 1 · 摘自论文原文
  • 从感知、通信到推理三层构建攻击框架
  • 在GQA上实现最高78.3%的攻击成功率
  • 适合研究系统安全与鲁棒性的人参考

多模态多智能体系统(MM-MAS)因其在跨模态复杂推理与协作方面的潜力而备受关注。随着系统规模和功能的扩展,其潜在漏洞也日益重要。然而,现有对抗攻击研究主要聚焦于孤立智能体或单模态场景,对MM-MAS的脆弱性仍缺乏探索。为此,我们提出HAM³——一种针对多模态多智能体系统的分层攻击框架,将攻击分解为三个相互关联的层级:在感知层,通过扰动视觉输入、文本输入及其融合表示实施攻击;在通信层,干扰消息内容与交互拓扑,如篡改共享上下文或通信链路以扭曲信息流;在推理层,干预智能体的认知流程,误导推理路径并最终影响决策。我们在基于ReAct、Plan-and-Solve和Reflexion等不同推理范式的多智能体系统上,于GQA基准上评估了HAM³。实验表明,该框架最高可达78.3%的攻击成功率,其中推理层攻击最为有效;超过一半的成功攻击导致多个智能体产生一致错误。这些发现为构建更鲁棒、可解释的多智能体智能提供了重要启示。

原文摘要 · Abstract (English)

Multi-modal multi-agent systems (MM-MAS) have gained increasing attention for their capacity to enable complex reasoning and coordination across diverse modalities. As these systems continue to expand in scale and functionality, investigating their potential vulnerabilities has become increasingly important. However, existing studies on adversarial attacks in multi-agent systems primarily focus on isolated agents or unimodal settings, leaving the vulnerabilities of MM-MAS largely underexplored. To bridge this gap, we introduce HAM$^{3}$, a Hierarchical Attack framework for multi-modal multi-agent systems that decomposes attacks into three interconnected layers. Specifically, at the perception layer, HAM$^{3}$ mounts attacks by perturbing visual inputs, textual inputs, and their fused visual-textual representations. At the communication layer, it performs communication-level attacks that corrupt message content and interaction topology, such as manipulating shared context or communication links to distort collective information flow. At the reasoning layer, it conducts reasoning-level attacks that interfere with each agent's cognitive pipeline, biasing reasoning trajectories and ultimately compromising final decisions. We evaluate HAM$^{3}$ on the GQA benchmark through multi-agent systems built on distinct reasoning paradigms including ReAct, Plan-and-Solve, and Reflexion. Experiments demonstrate that our framework achieves an Attack Success Rate of up to 78.3%, with reasoning-layer attacks being the most effective. More than half of the successful attacks lead multiple agents to produce consistent errors. These findings offer valuable insights for building more robust and interpretable multi-agent intelligence.

多智能体对抗攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。