arXiv:2608.24275cs.AIcs.CL2026-08

用强化学习让智能体自动判断何时启用安全策略。

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

论文配图:RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
图 1 · 摘自论文原文
  • 通过强化学习动态选择合适的安全策略
  • 在6个基准上实现强安全检测性能
  • 适合需要自适应安全防护的智能体系统

保障语言模型智能体的安全性需在上下文相关的安全策略下评估完整执行轨迹。现有政策感知防护主要依赖提示或监督微调,难以适应未见轨迹和变化的政策环境。我们提出RePolicy,一种通过强化学习学习安全策略调用的智能体防护机制。给定智能体轨迹和动态策略库,RePolicy识别适用策略,利用其内容生成基于策略的理由和安全判断。我们构建PolicyTraj-20K用于监督初始化,随后采用可验证奖励与策略上下文扰动的GRPO算法。在六个智能体安全基准上的实验表明,RePolicy在不同政策上下文中均实现优异的整体安全检测性能和稳健的策略调用能力。

原文摘要 · Abstract (English)

Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an agent safeguard that learns safety-policy invocation through reinforcement learning. Given an agent trajectory and a dynamic policy library, RePolicy invokes the applicable policy and uses its content to produce a policy-grounded rationale and safety judgment. We construct PolicyTraj-20K to support supervised initialization, followed by GRPO with verifiable rewards and policy-context perturbation. Experiments across six agent safety benchmarks show that RePolicy achieves strong overall safety-detection performance and robust policy invocation under varying policy contexts.

智能体安全强化学习策略调用动态防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。