提出动态安全罩,让机器人导航强化学习更安全高效。
A Dynamic Safety Shield for Safe and Efficient Reinforcement Learning of Navigation Tasks
- 用动态调整的控制策略融合强化学习与优化控制器
- 仿真中达成目标数更多、碰撞次数更少,优于现有方法
- 适合需高安全性和强探索能力的机器人导航场景
强化学习(RL)在机器人应用中表现优异,但其安全性及现实世界迁移仍是挑战。传统安全强化学习或因早期频繁碰撞,或因硬性约束抑制探索而效果受限。本文提出一种新型动态安全罩,结合优化控制器的鲁棒性与强化学习的长期预测能力,使RL代理能自适应调节控制器参数。该方法显著提升导航任务中的探索效率,同时大幅减少碰撞。仿真结果表明,在多种复杂环境中,本方法在“达成目标数/碰撞数”指标上优于现有先进基线;相比经典安全罩,达成目标更多;相比约束型RL,碰撞更少。最后,我们在真实机器人上验证了该方法的有效性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has been successfully applied to a variety of robotics applications, where it outperforms classical methods. However, the safety aspect of RL and the transfer to the real world remain an open challenge. A prominent field for tackling this challenge and ensuring the safety of the agents during training and execution is safe reinforcement learning. Safe RL can be achieved through constrained RL and safe exploration approaches. The former learns the safety constraints over the course of training to achieve a safe behavior by the end of training, at the cost of high number of collisions at earlier stages of the training. The latter offers robust safety by enforcing the safety constraints as hard constraints, which prevents collisions but hinders the exploration of the RL agent, resulting in lower rewards and poor performance. To overcome those drawbacks, we propose a novel safety shield, that combines the robustness of the optimization-based controllers with the long prediction capabilities of the RL agents, allowing the RL agent to adaptively tune the parameters of the controller. Our approach is able to improve the exploration of the RL agents for navigation tasks, while minimizing the number of collisions. Experiments in simulation show that our approach outperforms state-of-the-art baselines in the reached goals-to-collisions ratio in different challenging environments. The goals-to-collisions ratio metrics emphasizes the importance of minimizing the number of collisions, while learning to accomplish the task. Our approach achieves a higher number of reached goals compared to the classic safety shields and fewer collisions compared to constrained RL approaches. Finally, we demonstrate the performance of the proposed method in a real-world experiment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。