用分层强化学习提升自动驾驶决策通用性与效率
Scenario-Based Hierarchical Reinforcement Learning for Automated Driving Decision Making
- 分层策略:高层选动作模板,底层执行控制
- 在复杂场景中训练后仍能安全高效决策
- 适合研究自动驾驶决策与强化学习融合的学者
高度自动化驾驶系统的决策算法开发仍具挑战性,因系统需在开放复杂环境中安全运行。强化学习(RL)可直接从经验中学习完整决策策略,在简单驾驶任务中已展现潜力。然而,现有方法在复杂任务中缺乏泛化能力且学习效率低。为此,我们提出首个基于场景的分层强化学习框架SAD-RL,将高层策略与低层控制逻辑结合,高层选择动作模板,由底层逻辑评估并执行。场景化环境可精准控制训练体验,并显式引入高难度但稀有情境以增强鲁棒性。实验表明,使用SAD-RL训练的智能体可在简单与复杂场景中均实现安全行为,且学习高效。消融实验确认分层结构与场景多样性对性能至关重要。
原文摘要 · Abstract (English)
Developing decision-making algorithms for highly automated driving systems remains challenging, since these systems have to operate safely in an open and complex environments. Reinforcement Learning (RL) approaches can learn comprehensive decision policies directly from experience and already show promising results in simple driving tasks. However, current approaches fail to achieve generalizability for more complex driving tasks and lack learning efficiency. Therefore, we present Scenario-based Automated Driving Reinforcement Learning (SAD-RL), the first framework that integrates Reinforcement Learning (RL) of hierarchical policy in a scenario-based environment. A high-level policy selects maneuver templates that are evaluated and executed by a low-level control logic. The scenario-based environment allows to control the training experience for the agent and to explicitly introduce challenging, but rate situations into the training process. Our experiments show that an agent trained using the SAD-RL framework can achieve safe behaviour in easy as well as challenging situations efficiently. Our ablation studies confirmed that both HRL and scenario diversity are essential for achieving these results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。