用扰动观测器提升安全强化学习,允许模型不精确仍保持高安全性。
Disturbance Observer-based Control Barrier Functions with Residual Model Learning for Safe Reinforcement Learning
- 结合扰动观测器与残差建模,实现弱依赖精确动力学的安全部署
- 在Safety-gym基准上优于仅用残差模型或观测器的方法
- 适用于对安全要求高的物理系统,如真实机器人赛车
强化学习代理需探索环境以学习最优行为并获得最大奖励,但直接在真实系统上训练探索可能带来风险;而基于仿真的训练又存在仿真到现实的差距问题。近期方法利用安全过滤器(如控制屏障函数,CBF)在训练中惩罚不安全动作。然而,CBF的强安全保证依赖于精确的动力学模型。实际中,总存在不确定性,包括动态误差带来的内部扰动和风等外部扰动。本文提出一种基于扰动抑制的安全强化学习新框架,允许近似无模型的强化学习,仅需一个假设性的但不一定精确的名义动力学模型。我们在Safety-gym基准上的Point和Car机器人所有任务中验证了该方法,结果优于仅使用残差建模或扰动观测器的现有先进方法。进一步通过真实F1/10竞速车验证了该框架的有效性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) agents need to explore their environment to learn optimal behaviors and achieve maximum rewards. However, exploration can be risky when training RL directly on real systems, while simulation-based training introduces the tricky issue of the sim-to-real gap. Recent approaches have leveraged safety filters, such as control barrier functions (CBFs), to penalize unsafe actions during RL training. However, the strong safety guarantees of CBFs rely on a precise dynamic model. In practice, uncertainties always exist, including internal disturbances from the errors of dynamics and external disturbances such as wind. In this work, we propose a new safe RL framework based on disturbance rejection-guarded learning, which allows for an almost model-free RL with an assumed but not necessarily precise nominal dynamic model. We demonstrate our results on the Safety-gym benchmark for Point and Car robots on all tasks where we can outperform state-of-the-art approaches that use only residual model learning or a disturbance observer (DOB). We further validate the efficacy of our framework using a physical F1/10 racing car. Videos: https://sites.google.com/view/res-dob-cbf-rl
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。