分层强化学习提升复杂路况自动驾驶探索能力
Extensive Exploration in Complex Traffic Scenarios using Hierarchical Reinforcement Learning
- 分两阶段训练高层与底层控制器,解耦长期与短期决策
- 在高速复杂场景下实现更优的导航表现,适应延迟奖励
- 适合研究自动驾驶决策系统或强化学习应用者阅读
开发能够应对复杂交通环境的自动驾驶系统仍是重大挑战。与基于规则或监督学习的方法不同,基于深度强化学习(DRL)的控制器无需领域知识和数据集,具备更强适应性。然而,现有DRL研究多聚焦于简单交通模式,难以有效处理具有延迟、长期奖励的复杂驾驶环境,影响结果泛化性。为此,本文提出一种开创性的分层框架,将复杂的决策问题高效分解为可管理且可解释的子任务。采用两步训练流程,分别训练高层控制器与底层控制器:高层控制器利用长期延迟奖励增强探索能力,底层控制器通过短期即时奖励实现纵向与横向控制。仿真实验表明,该分层控制器在复杂高速公路驾驶场景中表现更优。
原文摘要 · Abstract (English)
Developing an automated driving system capable of navigating complex traffic environments remains a formidable challenge. Unlike rule-based or supervised learning-based methods, Deep Reinforcement Learning (DRL) based controllers eliminate the need for domain-specific knowledge and datasets, thus providing adaptability to various scenarios. Nonetheless, a common limitation of existing studies on DRL-based controllers is their focus on driving scenarios with simple traffic patterns, which hinders their capability to effectively handle complex driving environments with delayed, long-term rewards, thus compromising the generalizability of their findings. In response to these limitations, our research introduces a pioneering hierarchical framework that efficiently decomposes intricate decision-making problems into manageable and interpretable subtasks. We adopt a two step training process that trains the high-level controller and low-level controller separately. The high-level controller exhibits an enhanced exploration potential with long-term delayed rewards, and the low-level controller provides longitudinal and lateral control ability using short-term instantaneous rewards. Through simulation experiments, we demonstrate the superiority of our hierarchical controller in managing complex highway driving situations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。