arXiv:2602.05311cs.LGcs.AI2026-02

让神经控制器在动态扰动下仍能保证安全与稳定。

Formal Synthesis of Certifiably Robust Neural Lyapunov-Barrier Certificates

  • 基于Lipschitz连续性设计鲁棒神经李雅普诺夫-屏障函数
  • 在倒立摆和2D对接任务中,鲁棒性提升最高达4.6倍
  • 适合对安全性要求高的强化学习系统设计

神经李雅普诺夫和屏障函数最近被用作验证深度强化学习(RL)控制器安全性和稳定性的重要工具。然而,现有方法仅在理想无扰动动力学下提供保证,限制了其在存在不确定性的真实场景中的可靠性。本文研究了在系统动力学扰动下合成鲁棒神经李雅普诺夫-屏障函数的问题。我们形式化定义了鲁棒李雅普诺夫-屏障函数,并基于Lipschitz连续性提出了确保对有界扰动鲁棒性的充分条件。我们提出了实际的训练目标,通过对抗训练、Lipschitz邻域约束和全局Lipschitz正则化来强制满足这些条件。我们在两个具有实际意义的环境中验证了该方法:倒立摆(广泛研究的基准)和2D对接(自主系统中的关键安全任务)。结果表明,相比基线方法,我们的方法在强扰动下显著提升了认证鲁棒性边界(最高达4.6倍)和经验成功率(最高达2.4倍)。实验验证了在动态扰动下训练鲁棒神经证书的有效性。

原文摘要 · Abstract (English)

Neural Lyapunov and barrier certificates have recently been used as powerful tools for verifying the safety and stability properties of deep reinforcement learning (RL) controllers. However, existing methods offer guarantees only under fixed ideal unperturbed dynamics, limiting their reliability in real-world applications where dynamics may deviate due to uncertainties. In this work, we study the problem of synthesizing \emph{robust neural Lyapunov barrier certificates} that maintain their guarantees under perturbations in system dynamics. We formally define a robust Lyapunov barrier function and specify sufficient conditions based on Lipschitz continuity that ensure robustness against bounded perturbations. We propose practical training objectives that enforce these conditions via adversarial training, Lipschitz neighborhood bound, and global Lipschitz regularization. We validate our approach in two practically relevant environments, Inverted Pendulum and 2D Docking. The former is a widely studied benchmark, while the latter is a safety-critical task in autonomous systems. We show that our methods significantly improve both certified robustness bounds (up to $4.6$ times) and empirical success rates under strong perturbations (up to $2.4$ times) compared to the baseline. Our results demonstrate effectiveness of training robust neural certificates for safe RL under perturbations in dynamics.

强化学习安全控制神经验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。