arXiv:2605.28214cs.CRcs.LG2026-05中稿 · EMNLP被引 2

隐空间攻击可隐藏在多智能体系统中,不依赖文本仍能破坏任务性能。

Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

论文配图:Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems
图 1 · 摘自论文原文
  • 通过隐状态干预触发攻击效应,无需使用对抗性文本。
  • 在跨智能体键值缓存传递时,任务性能下降显著超过局部状态干扰。
  • 提醒研究者关注不可见的隐空间风险,需超越文本检查的防护策略。

基于隐空间的多智能体系统将部分显式通信替换为隐藏表示,提升了协作效率与灵活性。然而,将协调机制置于隐空间的同时,也可能使攻击行为脱离可见文本的检测范围。本文研究隐状态是否可在清洁执行中携带攻击相关信息并保持有效性。为此,我们提出一种隐空间攻击框架,通过隐状态干预重新激活攻击影响,而无需重复使用对抗性文本。大量实验表明,此类隐空间攻击在清洁执行中显著降低任务性能,尤其在跨智能体键值缓存传递(inter-agent KV-cache handoffs)中表现更明显。进一步控制分析显示,这种性能退化无法归因于任意扰动或无效生成。结果表明,隐空间协作并未消除攻击风险,而是将部分风险转移至更难观测的执行状态,亟需超越可见文本检查的防御机制。

原文摘要 · Abstract (English)

Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficient and flexible agent collaboration. However, moving coordination into latent space may also move attacks beyond the reach of visible-text inspection. In this paper, we study whether latent states can carry attack-associated information that remains effective during clean executions. To examine this question, we introduce a latent attack framework that reactivates attack-induced effects through latent interventions without reusing adversarial text. Extensive experiments show that the resulting latent attacks can substantially degrade task performance in clean executions, especially when applied to inter-agent KV-cache handoffs rather than local hidden states. Further control analyses indicate that this degradation cannot be reduced to arbitrary perturbations or invalid generation. Overall, our findings suggest that latent-based collaboration does not remove attack risk. It shifts part of the risk into less observable execution states, calling for safeguards beyond visible-text inspection.

多智能体隐空间攻击安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。