arXiv:2510.18082cs.LGcs.RO2025-10中稿 · publication in the…被引 2

安全过滤不降性能,理论证明可实现绝对安全与最优表现兼顾

Provably Optimal Reinforcement Learning under Safety Filtering

  • 用安全关键MDP建模,要求完全避免灾难状态
  • 最宽松的安全过滤器下仍能达成最优长期回报
  • 适合需绝对安全的机器人、自动驾驶等场景

强化学习在复杂任务中取得进展,但缺乏形式化安全保证限制其在关键领域的应用。常用做法是在策略上叠加安全过滤器以阻止危险动作,但常被认为会牺牲性能。本文首次证明:只要安全过滤器足够宽松,安全约束不会损害渐近性能。通过定义安全关键马尔可夫决策过程(SC-MDP),要求对灾难性状态进行类别化规避(非高概率规避)。进一步构建过滤后MDP,其中所有动作均安全,因安全过滤器被视为环境一部分。主定理表明:(i) 在过滤后MDP中学习是类别化安全的;(ii) 标准RL收敛性仍成立;(iii) 过滤后MDP中的最优策略,在相同过滤器下执行时,能达到与SC-MDP中最优安全策略相同的渐近回报。该结果实现安全与性能解耦。在Safety Gymnasium上验证,训练中零违规,最终性能匹配或超越无过滤基线。

原文摘要 · Abstract (English)

Recent advances in reinforcement learning (RL) enable its use on increasingly complex tasks, but the lack of formal safety guarantees still limits its application in safety-critical settings. A common practical approach is to augment the RL policy with a safety filter that overrides unsafe actions to prevent failures during both training and deployment. However, safety filtering is often perceived as sacrificing performance and hindering the learning process. We show that this perceived safety-performance tradeoff is not inherent and prove, for the first time, that enforcing safety with a sufficiently permissive safety filter does not degrade asymptotic performance. We formalize RL safety with a safety-critical Markov decision process (SC-MDP), which requires categorical, rather than high-probability, avoidance of catastrophic failure states. Additionally, we define an associated filtered MDP in which all actions result in safe effects, thanks to a safety filter that is considered to be a part of the environment. Our main theorem establishes that (i) learning in the filtered MDP is safe categorically, (ii) standard RL convergence carries over to the filtered MDP, and (iii) any policy that is optimal in the filtered MDP-when executed through the same filter-achieves the same asymptotic return as the best safe policy in the SC-MDP, yielding a complete separation between safety enforcement and performance optimization. We validate the theory on Safety Gymnasium with representative tasks and constraints, observing zero violations during training and final performance matching or exceeding unfiltered baselines. Together, these results shed light on a long-standing question in safety-filtered learning and provide a simple, principled recipe for safe RL: train and deploy RL policies with the most permissive safety filter that is available.

强化学习安全控制理论证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。