arXiv:2507.04049cs.CVcs.RO2025-07TPAMI被引 21

用强化扩散生成多样驾驶轨迹,突破端到端自动驾驶的模仿瓶颈

DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving

  • 结合强化学习与扩散模型生成多条合理轨迹
  • 在NAVSIM和nuScenes上实现显著多样性提升
  • 适合需要多样化决策的自动驾驶研究者

当前多数端到端自动驾驶方法依赖单一专家示范进行模仿学习,导致行为保守且同质化,难以适应复杂真实场景。本文提出DIVER框架,将强化学习与基于扩散的生成机制融合,生成多样且可行的行驶轨迹。核心在于:首先,模型基于地图元素和周围交通参与者,从单条真实轨迹生成多条参考轨迹,缓解仅依赖单一示范带来的局限;其次,利用强化学习引导扩散过程,通过奖励信号约束轨迹的安全性与多样性,提升其实际适用性与泛化能力。此外,针对传统L2度量无法有效捕捉多模态预测多样性的缺陷,提出新型多样性评估指标。在闭环NAVSIM与Bench2Drive基准,以及开环nuScenes数据集上的大量实验表明,DIVER显著提升了轨迹多样性,有效解决了模仿学习固有的模式崩溃问题。

原文摘要 · Abstract (English)

Most end-to-end autonomous driving methods rely on imitation learning from single expert demonstrations, often leading to conservative and homogeneous behaviors that limit generalization in complex real-world scenarios. In this work, we propose DIVER, an end-to-end driving framework that integrates reinforcement learning with diffusion-based generation to produce diverse and feasible trajectories. At the core of DIVER lies a reinforced diffusion-based generation mechanism. First, the model conditions on map elements and surrounding agents to generate multiple reference trajectories from a single ground-truth trajectory, alleviating the limitations of imitation learning that arise from relying solely on single expert demonstrations. Second, reinforcement learning is employed to guide the diffusion process, where reward-based supervision enforces safety and diversity constraints on the generated trajectories, thereby enhancing their practicality and generalization capability. Furthermore, to address the limitations of L2-based open-loop metrics in capturing trajectory diversity, we propose a novel Diversity metric to evaluate the diversity of multi-mode predictions.Extensive experiments on the closed-loop NAVSIM and Bench2Drive benchmarks, as well as the open-loop nuScenes dataset, demonstrate that DIVER significantly improves trajectory diversity, effectively addressing the mode collapse problem inherent in imitation learning.

自动驾驶扩散模型强化学习轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。