arXiv:2605.06992cs.LGstat.ML2026-05

安全泛化难,因安全本身比执行更复杂。

Why Does Agentic Safety Fail to Generalize Across Tasks?

  • 安全约束使任务到控制器的映射更敏感,泛化难度上升。
  • 实验证明神经网络与大模型在新任务中安全表现退化。
  • 适合关注AI安全泛化能力的研究者阅读。

AI代理在多任务场景中日益普及,任务在测试时动态指定,代理需泛化至未见任务。核心关切是安全性:代理不仅需执行新任务,还需规避风险并处理突发情况。经验表明,尽管执行能力可泛化,安全能力却常失效。本文通过理论与实验指出,这种失败并非训练方法局限,而是安全本质属性所致——任务与安全执行的关系远比任务与执行本身复杂。理论上,分析带 $H_{ u}$-鲁棒性的线性二次控制,证明加入安全要求后,任务到最优控制器的映射具有更高Lipschitz常数,给出一个独立有趣的界。实证上,在神经网络代理的四旋翼导航与大模型代理的客户关系管理(CRM)中均验证结论。研究暗示当前提升代理安全的努力可能不足,需根本性新方法。

原文摘要 · Abstract (English)

AI agents are increasingly deployed in multi-task settings, where the task to perform is specified at test time, and the agent must generalize to unseen tasks. A major concern in such settings is safety: often, an agent must not only execute unseen tasks, but do so while avoiding risks and handling ones that materialize. Empirical evidence suggests that even when the ability to execute generalizes to unseen tasks, the ability to do so safely frequently does not. This paper provides theory and experiments indicating that failures of agentic safety to generalize across tasks are not merely due to limitations of training methods, but reflect an inherent property of safety itself: the relationship between a task and its safe execution is more complex than the relationship between a task and its execution alone. Theoretically, we analyze linear-quadratic control with $H_{\infty}$-robustness, and prove that the mapping from task specification to an optimal controller has higher Lipschitz constant with safety requirements than without, yielding a Lipschitz bound of independent interest. Empirically, we demonstrate our conclusions in simulated quadcopter navigation with a neural network agent and in CRM with an LLM agent. Our findings suggest that current efforts to enhance agentic safety may be insufficient, and point to a need for fundamentally different approaches.

AI安全泛化能力强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。