arXiv:2601.21297cs.ROcs.SY2026-01中稿 · the 8th Annual Lea…

无需模型即可学习安全约束,让智能体更安全高效地学习控制。

Deep QP Safety Filter: Model-free Learning for Reachability-based Safety Filter

  • 用无模型学习结合可达性理论构建安全过滤器
  • 在多种系统上显著减少学习初期失败次数
  • 适合需要安全性的强化学习场景

我们提出 Deep QP Safety Filter,一种完全数据驱动的安全层,用于黑箱动态系统。该方法通过结合哈密顿-雅可比(HJ)可达性与无模型学习,在无需系统模型的情况下学习二次规划(QP)安全过滤器。我们为安全值及其导数构建基于收缩的损失函数,并相应训练两个神经网络。在理想情况下,学习到的评价器收敛至粘性解(及其导数),即使对于非光滑值也成立。在多种动态系统——包括混合系统——及多个强化学习任务中,Deep QP Safety Filter 显著减少了预收敛阶段的失败次数,同时加速了向更高回报的学习过程,提供了一条安全且实用的无模型控制路径。

原文摘要 · Abstract (English)

We introduce Deep QP Safety Filter, a fully data-driven safety layer for black-box dynamical systems. Our method learns a Quadratic-Program (QP) safety filter without model knowledge by combining Hamilton-Jacobi (HJ) reachability with model-free learning. We construct contraction-based losses for both the safety value and its derivatives, and train two neural networks accordingly. In the exact setting, the learned critic converges to the viscosity solution (and its derivative), even for non-smooth values. Across diverse dynamical systems -- even including a hybrid system -- and multiple RL tasks, Deep QP Safety Filter substantially reduces pre-convergence failures while accelerating learning toward higher returns than strong baselines, offering a principled and practical route to safe, model-free control.

安全控制强化学习无模型学习可达性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。