arXiv:2605.28201cs.AI2026-05被引 1

攻击者植入恶意内容,潜伏多轮交互后突然触发危险行为。

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

论文配图:Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
图 1 · 摘自论文原文
  • 恶意内容藏于代理状态中,跨轮次潜伏不被察觉。
  • 实验发现7个主流模型均易受此攻击,即使单轮攻击成功率低。
  • 适合关注AI安全、代理系统防护的研究者与开发者。

大型语言模型(LLM)代理仍易受外部环境的安全威胁,攻击者可通过工具返回数据、网页或MCP上下文等外部观测注入对抗性内容,引发代理产生不安全行为或错误输出。现有研究多聚焦单轮交互攻击,即代理在一次请求中立即表现出危害行为。本文揭示,对抗性内容可长期存在于代理状态中,跨多轮交互保持潜伏,并在后续正常用户查询下被激活。我们将其定义为「潜伏攻击」(Sleeper Attack)。为评估该威胁,构建包含1,896个实例的基准测试,涵盖六种真实世界有害结果、三种攻击策略和三个代理状态目标:会话上下文、记忆和可复用技能。在七个主流开源与闭源大模型上进行实验,结果显示当前最先进的代理系统仍对潜伏攻击高度脆弱,即便在单轮攻击基线中表现良好。代码与数据已公开。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP context, causing harmful agentic behaviors such as unsafe actions or incorrect outputs. Existing studies typically focus on single-interaction attacks, where the agent observes adversarial content and immediately exhibits harmful behavior within one user request. However, we show that adversarial content can also persist across interactions served by the same agent, making such threats harder to detect and mitigate. Specifically, adversarial content may persist in the agent state, remain dormant across interactions, and later be activated by a benign user query. We formalize this type of safety threat as Sleeper Attack. To evaluate it, we construct a benchmark with 1,896 instances covering six real-world harmful outcomes, three attack strategies, and three agent state targets: session context, memory, and reusable skills. Experiments on seven strong open-source and closed-source LLMs show that state-of-the-art LLM agents remain vulnerable to Sleeper Attack, even when they achieve low attack success rates under a single-interaction baseline. Our code and data are available at https://anonymous.4open.science/r/skdvnfu23ihr9wdscnksf1asdffsaef.

AI安全模型攻击语言模型代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。