提出安全导航新算法,让机器人在不依赖模型情况下实现高效且可验证的自主避障。
Certificated Actor-Critic: Hierarchical Reinforcement Learning with Control Barrier Functions for Safe Navigation
- 分层强化学习框架结合控制屏障函数设计安全奖励机制。
- 理论证明算法可保证导航安全性,仿真中成功规避所有障碍物。
- 适合对安全性和可靠性要求高的机器人自主系统研发者。
控制屏障函数(CBFs)已成为设计机器人安全导航系统的重要方法。然而现有基于CBF的方法存在局限:基于优化的安全控制通常具有短视性或计算成本高,且依赖简化系统模型;而基于学习的方法则缺乏导航性能与安全性的定量指标。本文提出一种新型无模型强化学习算法——认证演员-评论家(Certificated Actor-Critic, CAC),构建了分层强化学习框架,并基于CBFs设计了明确的奖励函数。我们进行了理论分析与算法证明,并提出了多项实现改进。通过两个仿真实验验证了所提CAC算法的有效性,结果表明该方法在无需系统模型的前提下实现了高效且可证明的安全导航。
原文摘要 · Abstract (English)
Control Barrier Functions (CBFs) have emerged as a prominent approach to designing safe navigation systems of robots. Despite their popularity, current CBF-based methods exhibit some limitations: optimization-based safe control techniques tend to be either myopic or computationally intensive, and they rely on simplified system models; conversely, the learning-based methods suffer from the lack of quantitative indication in terms of navigation performance and safety. In this paper, we present a new model-free reinforcement learning algorithm called Certificated Actor-Critic (CAC), which introduces a hierarchical reinforcement learning framework and well-defined reward functions derived from CBFs. We carry out theoretical analysis and proof of our algorithm, and propose several improvements in algorithm implementation. Our analysis is validated by two simulation experiments, showing the effectiveness of our proposed CAC algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。