对抗性提示注入攻击与防御的动态演化框架,提升系统安全性。
CoRL: Co-Evolutionary Reinforcement Learning for Adaptive Indirect Prompt-Injection Attacks and Defenses

- 构建攻防双角色博弈模型,动态调整攻击策略和防御机制。
- 在1514次测试中将攻击成功率降至0.0%,任务完成度提升至76.3%。
- 适用于安全评估、红队测试及持续演化的智能系统防御设计。
工具增强型语言代理易受间接提示注入(IPI)攻击。与直接注入不同,IPI将恶意指令隐藏在不可信工具输出中,隐蔽改变合法任务执行流程。传统防御因固定攻击模式而失效,当攻击者改变策略、注入位置或载荷时,防御能力下降。为此,本文将自适应IPI建模为不对称、部分可观测、通用和博弈:多轮攻击者根据公开轨迹中的工具返回点动态调整载荷,而工具使用者防御者需阻止注入目标并完成用户任务。提出CoRL框架,包含三个阶段:攻击者SFT基于成功轨迹初始化多轮攻击;双边Co-PPO联合训练双智能体,采用角色特异性奖励与历史对手种群;防御者SFT整合验证器接受的教师修复方案,固化种群发现的缺陷修复。在每名防御者1,514次干净、固定模板及自适应执行中,CoRL将总体攻击成功率(ASR)降低38.5个百分点至0.0%,任务效用提升13.1个百分点至76.3%。阶段与控制消融实验表明在线Co-PPO与种群挖掘修复均有正向贡献,外部基准评估显示攻击抵抗能力具有迁移性。防御者在评估攻击下平衡安全性与任务效用,保留的攻击者可作为自适应红队评估候选。
原文摘要 · Abstract (English)
Tool-augmented language agents are vulnerable to indirect prompt injection (IPI). Unlike direct prompt injection, IPI hides adversarial instructions in untrusted tool outputs and can covertly alter the execution of a legitimate task. Defenses trained on fixed attacks may fail as an attacker changes its strategy, injection site, and payload. To address this problem, we formulate adaptive IPI as an asymmetric, partially observable, general-sum Markov game: a multi-turn attacker adapts payloads at reached tool-return sites from the public trajectory, while a tool-using defender must block the injected objective and complete the user task. We propose CoRL, a verifier-grounded co-evolution and repair framework with three stages: Attacker SFT initializes multi-turn attacks from successful trajectories; bilateral Co-PPO jointly trains both agents with role-specific rewards and historical opponent populations; and Defender SFT consolidates verifier-accepted teacher repairs for population-discovered failures. Across 1,514 clean, fixed-template, and adaptive executions per defender, CoRL reduces overall ASR by 38.5 points to 0.0% and raises utility by 13.1 points to 76.3%. Stage-wise and controlled ablations show positive contributions from online Co-PPO and population-mined repair, while external-benchmark evaluation indicates transfer in attack resistance. The defender balances safety and task utility under the evaluated attacks, while the retained attackers provide candidates for adaptive red-team evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。