arXiv:2607.19913cs.AIcs.CL2026-07

提前预见长周期智能体风险,防止工具使用中出现安全问题

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

论文配图:JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
图 1 · 摘自论文原文
  • 通过多智能体仿真生成轨迹,训练守卫模型预测潜在风险
  • 在4个基准上提升安全防护15.9个百分点,同时提升任务完成率5.1个百分点
  • 适合关注长周期智能体安全与风险预判的研究者和开发者

智能体安全正从内容审核转向预防工具使用前的操作失败。我们提出Janus,一种面向长周期智能体安全的前瞻性框架,训练守卫模型从部分轨迹中预见延迟风险。Janus通过多智能体仿真合成多样化轨迹,学习一个共享策略,包含两个耦合任务:预测安全相关未来(预期任务)和基于已观测前缀与预期未来的安全判断(裁决任务)。两项任务通过CoAA-RL联合优化,以预测对下游安全判断的实用性为奖励。最终得到的守卫模型Vanguard可在动作执行前拦截不安全行为。在四个智能体安全基准测试中,Vanguard相比基线守卫平均保护能力提升15.9个百分点,同时良性任务完成率提升5.1个百分点。

原文摘要 · Abstract (English)

Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution. Across four agent-safety benchmarks, Vanguard improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.

智能体安全风险预测强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。