arXiv:2604.13192eess.SYcs.RO2026-04

用强化学习方法在未知系统上合成更安全的鲁棒控制屏障函数。

Synthesis and Deployment of Maximal Robust Control Barrier Functions through Adversarial Reinforcement Learning

论文配图:Synthesis and Deployment of Maximal Robust Control Barrier Functions through Adversarial Reinforcement Learning
图 1 · 摘自论文原文
  • 将安全值函数作为鲁棒控制屏障,直接从动态规划中推导。
  • 引入Q函数机制,无需显式动力学模型,提升通用性。
  • 在倒立摆和四足机器人上验证,比传统方法更少保守、更可靠。

鲁棒控制屏障函数(CBFs)为最坏情况扰动下的安全约束提供了系统化机制。然而,现有方法通常依赖于动力学中的显式、封闭形式结构(如控制仿射)和不确定性模型,导致可扩展性和通用性受限,大多数鲁棒CBFs仅能保证保守的极大鲁棒安全集子集。本文提出一种适用于一般非线性系统在有界不确定性下的新鲁棒CBF框架。首先证明求解动态规划Isaacs方程的安全值函数是有效的鲁棒离散时间CBF,可在极大鲁棒安全集上强制执行安全。随后引入强化学习中的质量函数(或Q函数)概念,通过将屏障证书提升至状态-动作空间,摆脱对显式动力学的需求,提出一种新型鲁棒Q-CBF安全过滤约束。结合对抗性强化学习,该方法可在黑箱动力学和未知不确定性结构的一般非线性系统上实现鲁棒Q-CBF的合成与部署。我们在典型的倒立摆基准测试和36维四足机器人模拟器上进行了验证,在倒立摆上获得显著更少保守的安全集,且在四足机器人上面对对抗性不确定性仍能可靠保证安全。

原文摘要 · Abstract (English)

Robust control barrier functions (CBFs) provide a principled mechanism for smooth safety enforcement under worst-case disturbances. However, existing approaches typically rely on explicit, closed-form structure in the dynamics (e.g., control-affine) and uncertainty models. This has led to limited scalability and generality, with most robust CBFs certifying only conservative subsets of the maximal robust safe set. In this paper, we introduce a new robust CBF framework for general nonlinear systems under bounded uncertainty. We first show that the safety value function solving the dynamic programming Isaacs equation is a valid robust discrete-time CBF that enforces safety on the maximal robust safe set. We then adopt the key reinforcement learning (RL) notion of quality function (or Q-function), which removes the need for explicit dynamics by lifting the barrier certificate into state-action space and yields a novel robust Q-CBF constraint for safety filtering. Combined with adversarial RL, this enables the synthesis and deployment of robust Q-CBFs on general nonlinear systems with black-box dynamics and unknown uncertainty structure. We validate the framework on a canonical inverted pendulum benchmark and a 36-D quadruped simulator, achieving substantially less conservative safe sets than barrier-based baselines on the pendulum and reliable safety enforcement even under adversarial uncertainty realizations on the quadruped.

控制屏障强化学习安全性非线性系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。