arXiv:2505.19516cs.RO2025-05被引 13

用扩散模型与监督策略融合,提升自动驾驶端到端控制的鲁棒性。

DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy

  • 混合扩散-监督解码器,结合生成能力与可控性。
  • 在CARLA和NAVSIM上达顶尖性能,长尾场景表现优异。
  • 适合研究自动驾驶、具身智能的学者与工程师。

端到端学习已成为自动驾驶的变革性范式,但驾驶行为的多模态特性和长尾场景下的泛化挑战仍是部署难题。本文提出DiffE2E,一种基于扩散模型的端到端自动驾驶框架。该框架首先通过分层双向交叉注意力机制对多传感器感知特征进行多尺度对齐;随后引入基于Transformer的新型混合扩散-监督解码器,并采用协同训练范式,无缝融合扩散与监督策略的优势。DiffE2E建模结构化潜在空间,其中扩散模型捕捉未来轨迹分布,监督部分增强可控性与鲁棒性。全局条件融合模块实现感知特征与高层目标的深度融合,显著提升轨迹生成质量。跨注意力机制促进融合特征与混合潜在变量间的高效交互,推动扩散与监督目标的联合优化,实现更稳健的输出生成。实验表明,DiffE2E在CARLA闭环评估与NAVSIM基准测试中均达到当前最优性能。所提出的融合策略为混合动作表示提供可推广范式,具备扩展至具身智能等更广领域的潜力。更多细节与可视化见项目网站。

原文摘要 · Abstract (English)

End-to-end learning has emerged as a transformative paradigm in autonomous driving. However, the inherently multimodal nature of driving behaviors and the generalization challenges in long-tail scenarios remain critical obstacles to robust deployment. We propose DiffE2E, a diffusion-based end-to-end autonomous driving framework. This framework first performs multi-scale alignment of multi-sensor perception features through a hierarchical bidirectional cross-attention mechanism. It then introduces a novel class of hybrid diffusion-supervision decoders based on the Transformer architecture, and adopts a collaborative training paradigm that seamlessly integrates the strengths of both diffusion and supervised policy. DiffE2E models structured latent spaces, where diffusion captures the distribution of future trajectories and supervision enhances controllability and robustness. A global condition integration module enables deep fusion of perception features with high-level targets, significantly improving the quality of trajectory generation. Subsequently, a cross-attention mechanism facilitates efficient interaction between integrated features and hybrid latent variables, promoting the joint optimization of diffusion and supervision objectives for structured output generation, ultimately leading to more robust control. Experiments demonstrate that DiffE2E achieves state-of-the-art performance in both CARLA closed-loop evaluations and NAVSIM benchmarks. The proposed integrated diffusion-supervision policy offers a generalizable paradigm for hybrid action representation, with strong potential for extension to broader domains including embodied intelligence. More details and visualizations are available at \href{https://infinidrive.github.io/DiffE2E/}{project website}.

自动驾驶扩散模型端到端强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。