用伪仿真强化学习提升端到端自动驾驶闭环表现
PerlAD: Towards Enhanced Closed-loop End-to-end Autonomous Driving with Pseudo-simulation-based Reinforcement Learning
- 基于离线数据构建向量空间的伪仿真环境,免渲染高效训练
- 预测世界模型生成动态响应轨迹,解决静态数据与动态驾驶的差距
- 分层解耦规划:模仿学习定横向路径,强化学习调纵向速度
基于模仿学习的端到端自动驾驶策略在闭环执行中常因训练目标与实际驾驶需求不匹配而表现不佳。虽然强化学习可通过奖励信号直接优化驾驶目标,但依赖渲染的训练环境存在渲染差距且计算成本高。为此,本文提出一种新型伪仿真强化学习方法PerlAD,基于离线数据构建运行于向量空间的伪仿真环境,实现无渲染、高效的试错训练。为弥合静态数据与动态闭环环境之间的鸿沟,PerlAD引入预测世界模型,根据自车规划生成响应式代理轨迹。此外,为提升规划效率,采用分层解耦规划器:模仿学习负责横向路径生成,强化学习优化纵向速度。大量实验表明,PerlAD在Bench2Drive基准上达到顶尖性能,驾驶得分超越前序端到端强化学习方法10.29%,且无需昂贵在线交互。在DOS基准上的额外评估也验证了其在安全关键遮挡场景下的可靠性。
原文摘要 · Abstract (English)
End-to-end autonomous driving policies based on Imitation Learning (IL) often struggle in closed-loop execution due to the misalignment between inadequate open-loop training objectives and real driving requirements. While Reinforcement Learning (RL) offers a solution by directly optimizing driving goals via reward signals, the rendering-based training environments introduce the rendering gap and are inefficient due to high computational costs. To overcome these challenges, we present a novel Pseudo-simulation-based RL method for closed-loop end-to-end autonomous driving, PerlAD. Based on offline datasets, PerlAD constructs a pseudo-simulation that operates in vector space, enabling efficient, rendering-free trial-and-error training. To bridge the gap between static datasets and dynamic closed-loop environments, PerlAD introduces a prediction world model that generates reactive agent trajectories conditioned on the ego vehicle's plan. Furthermore, to facilitate efficient planning, PerlAD utilizes a hierarchical decoupled planner that combines IL for lateral path generation and RL for longitudinal speed optimization. Comprehensive experimental results demonstrate that PerlAD achieves state-of-the-art performance on the Bench2Drive benchmark, surpassing the previous E2E RL method by 10.29% in Driving Score without requiring expensive online interactions. Additional evaluations on the DOS benchmark further confirm its reliability in handling safety-critical occlusion scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。