arXiv:2608.08795cs.CRcs.CL2026-08

让大模型攻击者用一次尝试就成功,靠的是提前学好策略。

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

论文配图:Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection
图 1 · 摘自论文原文
  • 攻击者通过分析成功失败轨迹,离线提炼通用攻击策略。
  • 测试时仅需一次交互,攻击成功率比之前最高提升11.8%。
  • 策略可跨防御机制迁移,适合研究模型安全的人员。

使用工具的大型语言模型代理易受间接提示注入(IPI)攻击,恶意指令嵌入外部观测值中可操控后续决策与行为。现有自适应攻击多依赖反复查询与调优,但真实攻击者可能仅有一次与未知目标代理交互的机会。本文提出SAVOR(基于结果条件反射的策略抽象),将攻击适应从测试阶段迭代转为离线策略提炼。SAVOR对来自独立训练环境的成功与失败轨迹进行结果条件反射,验证上下文相关的候选策略,并迭代整合为可复用的策略记忆。测试时,冻结的记忆指导生成针对每个未知目标的单次有效载荷,仅需一次目标代理查询且无需反馈。在两个基准和三个受害模型上,SAVOR在全部六组设置中均达到最高平均攻击成功率,较最强先前攻击提升2.5至11.8个百分点;在不提供攻击工具的Agent Security Bench上领先23.1点,在我们提出的可执行基准OpenClaw-IPI(通过工具交互与执行凭证验证攻击)上领先28.6点。在一种防御下学习的策略还能迁移至其他防御。

原文摘要 · Abstract (English)

Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external observations manipulate subsequent agent decisions and actions. Most existing adaptive attacks rely on repeatedly querying and refining against the target agent, whereas realistic attackers may have only a single opportunity to interact with an unknown target agent. We propose SAVOR (Strategy Abstraction Via Outcome-Conditioned Reflection), which shifts attack adaptation from test-time iteration to offline strategy distillation. SAVOR performs outcome-conditioned reflection over successful and failed trajectories collected from disjoint training environments, validates context-conditioned candidate strategies, and iteratively consolidates them into a reusable strategy memory. At test time, the frozen memory guides the generation of a single payload for each unseen target, requiring only one target-agent query and no target-agent feedback. Across two benchmarks and three victim models, SAVOR attains the highest average attack success rate in all six settings, leading the strongest prior attack by 2.5 to 11.8 points and the same injection channel without strategy learning by 23.1 points on Agent Security Bench, which holds out attacker tools, and 28.6 points on OpenClaw-IPI, an executable benchmark we introduce that holds out attack goals and verifies attacks through tool interactions and execution receipts. A memory learned under one defense also transfers to another.

提示注入攻击策略大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。