提出分层扩散与解耦强化学习,提升自动驾驶端到端规划性能
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
- 分层扩散生成轨迹,从粗到细逐步优化
- 生成轨迹保持运动结构,减少不现实路径
- 多目标解耦强化学习,训练更高效,适合自动驾驶系统
端到端规划已成为自动驾驶主流范式,现有模型多采用评分-选择框架从大量候选轨迹中选取最优路径,基于扩散模型的解码展现出强大潜力。然而,直接从完整候选空间中选择仍难优化,且扩散过程中的高斯扰动常引入不现实轨迹,增加去噪难度。此外,强化学习虽具前景,但现有端到端方法通常依赖单一耦合奖励,缺乏结构化信号,限制优化效果。为此,本文提出HAD框架,采用分层扩散策略,将规划分解为粗粒度到细粒度的过程;引入结构保持轨迹扩展方法,生成符合运动规律的候选轨迹;设计度量解耦策略优化(MDPO),实现多驾驶目标下的结构化强化学习。大量实验表明,HAD在NAVSIM和HUGSIM上均达到新基准:NAVSIM上提升2.3 EPDMS,HUGSIM上提升4.9路线完成率。
原文摘要 · Abstract (English)
End-to-end planning has emerged as a dominant paradigm for autonomous driving, where recent models often adopt a scoring-selection framework to choose trajectories from a large set of candidates, with diffusion-based decoding showing strong promise. However, directly selecting from the entire candidate space remains difficult to optimize, and Gaussian perturbations used in diffusion often introduce unrealistic trajectories that complicate the denoising process. In addition, for training these models, reinforcement learning (RL) has shown promise, but existing end-to-end RL approaches typically rely on a single coupled reward without structured signals, limiting optimization effectiveness. To address these challenges, we propose HAD, an end-to-end planning framework with a Hierarchical Diffusion Policy that decomposes planning into a coarse-to-fine process. To improve trajectory generation, we introduce Structure-Preserved Trajectory Expansion, which produces realistic candidates while maintaining kinematic structure. For policy learning, we develop Metric-Decoupled Policy Optimization (MDPO) to enable structured RL optimization across multiple driving objectives. Extensive experiments show that HAD achieves new state-of-the-art performance on both NAVSIM and HUGSIM, outperforming prior arts by a huge margin: +2.3 EPDMS on NAVSIM and +4.9 Route Completion on HUGSIM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。