分层强化学习让自动驾驶在路口更安全高效
Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections
- 分层结构结合目标引导的碰撞预测,提升决策安全性
- 训练速度更快,收敛到最优策略所需样本减少
- 适合复杂路口场景的自动驾驶系统研发者
强化学习在自动驾驶任务中展现巨大潜力,但在复杂场景下难以实现高效且安全的策略训练。本文提出一种新型分层强化学习框架,包含目标条件化的碰撞预测(GCCP)模块。在分层结构中,GCCP模块根据自车不同潜在子目标预测碰撞风险,高层决策器选择最安全的子目标,低层运动规划器据此与环境交互。相比传统方法,该算法更具样本效率,因子目标策略可在相似任务间复用。此外,GCCP模块能基于不同子目标预测自车及周边车辆的未来行为,确保决策全过程的安全性。实验表明,该方法收敛更快,安全性能优于传统强化学习方法。
原文摘要 · Abstract (English)
Reinforcement learning (RL) exhibits remarkable potential in addressing autonomous driving tasks. However, it is difficult to train a sample-efficient and safe policy in complex scenarios. In this article, we propose a novel hierarchical reinforcement learning (HRL) framework with a goal-conditioned collision prediction (GCCP) module. In the hierarchical structure, the GCCP module predicts collision risks according to different potential subgoals of the ego vehicle. A high-level decision-maker choose the best safe subgoal. A low-level motion-planner interacts with the environment according to the subgoal. Compared to traditional RL methods, our algorithm is more sample-efficient, since its hierarchical structure allows reusing the policies of subgoals across similar tasks for various navigation scenarios. In additional, the GCCP module's ability to predict both the ego vehicle's and surrounding vehicles' future actions according to different subgoals, ensures the safety of the ego vehicle throughout the decision-making process. Experimental results demonstrate that the proposed method converges to an optimal policy faster and achieves higher safety than traditional RL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。