用数据学习动态模型,实现机器人强化学习的零违规安全控制。
Robust Koopman Control Barrier Filters for Safe Actor-Critic Reinforcement Learning

- 基于数据构建有限维Koopman预测器,在提升空间中设计仿射安全约束。
- 通过二次规划安全层与残差余量修正,实现训练与部署中零约束违反。
- 适合对安全性要求高的机器人控制任务,尤其适用于无模型强化学习场景。
机器人系统的安全强化学习需在训练与部署阶段均满足状态和输入约束的同时提升任务性能。控制屏障函数(CBFs)可通过最小干预实现前向不变性,但其在无模型强化学习中的应用受限于对精确动力学的依赖及手动设计的屏障证书。本文提出鲁棒Koopman-CBF SAC:从数据中学习有限维Koopman预测器,在升维空间构造仿射CBF约束,并通过二次规划安全层实施。为应对有限维Koopman近似误差,采用保留轨迹数据估计的投影残差裕度收紧CBF条件。评论者在执行安全动作上训练,而演员则被正则化至Koopman-CBF可行集,减少对过滤器的依赖。在多个安全控制基准测试中,该方法在倒立摆稳定与跟踪任务中实现零约束违反,且性能匹配或超过无约束SAC。在高维Safety Gymnasium运动任务中,部分设置下减少违规,但也暴露一阶速度屏障与线性EDMD模型的局限性,推动高阶与多步Koopman-CBF扩展。结果表明,鲁棒Koopman-CBF滤波器是连接无模型强化学习与可验证安全性的有力桥梁,同时明确了此类滤波器有效运行的结构性条件。
原文摘要 · Abstract (English)
Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints during both training and deployment. Control barrier functions (CBFs) provide a principled mechanism for enforcing forward invariance through minimally invasive safety filters, but their use in model-free RL is limited by the need for accurate dynamics and hand-designed barrier certificates. We propose Robust Koopman-CBF SAC, a safety-filtered actor--critic framework that learns a finite-dimensional Koopman predictor from data, constructs affine CBF constraints in the lifted space, and enforces them through a quadratic-program safety layer. To account for finite-dimensional Koopman approximation error, the CBF condition is tightened using a projected residual margin estimated from held-out rollout data. The critic is trained on the executed safe action, while the actor is regularized toward the Koopman-CBF feasible set, reducing dependence on the filter over training. Across safe-control benchmarks, the method achieves zero constraint violations on CartPole stabilization and tracking while matching or exceeding unconstrained SAC returns. On high-dimensional Safety Gymnasium locomotion tasks, the method reduces violations in some settings but also exposes important limitations of first-order velocity barriers and linear EDMD models, motivating high-order and multi-step Koopman-CBF extensions. These results suggest that robust Koopman-CBF filters are a promising bridge between model-free RL and certifiable safety, while clarifying the structural conditions under which such filters remain effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。