arXiv:2504.00218cs.MAcs.AI2025-04ACL被引 43

针对多智能体LLM系统设计高效提示攻击,突破通信与安全限制。

$\textit{Agents Under Siege}$: Breaking Pragmatic Multi-Agent LLM Systems with Optimized Prompt Attacks

  • 基于最大流最小成本建模,优化提示分布以绕过分布式防御。
  • 在多种模型上攻击成功率提升最高达7倍,现有防御几乎失效。
  • 适合研究多智能体安全、对抗攻击或防御机制的学者参考。

当前大语言模型安全研究多集中于单智能体场景,但多智能体系统因依赖智能体间通信与去中心化推理,产生新型对抗风险。本文聚焦于受制于令牌带宽限制、消息延迟及防御机制的实用型多智能体系统,提出一种排列不变的对抗性攻击方法,通过优化在时延与带宽约束网络拓扑中的提示分布,突破系统内部分布式安全机制。将攻击路径建模为最大流最小成本问题,结合新颖的排列不变逃避损失(PIEL),利用图优化方法最大化攻击成功率同时最小化被检测风险。在Llama、Mistral、Gemma、DeepSeek等模型上,于JailBreakBench与AdversarialBench等数据集上评估,本方法相较传统攻击最高提升7倍攻击成功率,暴露出多智能体系统的严重漏洞。此外,包括Llama-Guard和PromptGuard在内的现有防御机制均无法阻止该攻击,凸显亟需专门面向多智能体场景的安全防护体系。

原文摘要 · Abstract (English)

Most discussions about Large Language Model (LLM) safety have focused on single-agent settings but multi-agent LLM systems now create novel adversarial risks because their behavior depends on communication between agents and decentralized reasoning. In this work, we innovatively focus on attacking pragmatic systems that have constrains such as limited token bandwidth, latency between message delivery, and defense mechanisms. We design a $\textit{permutation-invariant adversarial attack}$ that optimizes prompt distribution across latency and bandwidth-constraint network topologies to bypass distributed safety mechanisms within the system. Formulating the attack path as a problem of $\textit{maximum-flow minimum-cost}$, coupled with the novel $\textit{Permutation-Invariant Evasion Loss (PIEL)}$, we leverage graph-based optimization to maximize attack success rate while minimizing detection risk. Evaluating across models including $\texttt{Llama}$, $\texttt{Mistral}$, $\texttt{Gemma}$, $\texttt{DeepSeek}$ and other variants on various datasets like $\texttt{JailBreakBench}$ and $\texttt{AdversarialBench}$, our method outperforms conventional attacks by up to $7\times$, exposing critical vulnerabilities in multi-agent systems. Moreover, we demonstrate that existing defenses, including variants of $\texttt{Llama-Guard}$ and $\texttt{PromptGuard}$, fail to prohibit our attack, emphasizing the urgent need for multi-agent specific safety mechanisms.

多智能体对抗攻击提示安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。