arXiv:2503.12170cs.ROcs.CV2025-03被引 33

用扩散模型统一自动驾驶感知、预测与规划,简化系统结构。

DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving

  • 将自动驾驶建模为条件图像生成任务,统一多任务处理
  • 在Carla上实现新最高成功率与驾驶得分
  • 适合关注端到端系统简化与协同优化的研究者

端到端自动驾驶(E2E-AD)正成为实现完全自动化的有前景方法。然而,现有E2E-AD系统通常采用传统多任务框架,通过独立的任务专用头部处理感知、预测和规划任务。尽管以全可微方式训练,仍存在任务协调问题,系统复杂度高。本文提出DiffAD,一种新型扩散概率模型,将自动驾驶重新定义为条件图像生成任务。通过将异构目标栅格化至统一鸟瞰图(BEV),并建模其潜在分布,DiffAD在单一框架中统一多种驾驶目标,联合优化所有驾驶任务,显著降低系统复杂度并增强任务协调性。反向过程迭代精炼生成的BEV图像,带来更鲁棒、真实的驾驶行为。在Carla中的闭环评估表明,该方法取得新的最先进成功率与驾驶评分。

原文摘要 · Abstract (English)

End-to-end autonomous driving (E2E-AD) has rapidly emerged as a promising approach toward achieving full autonomy. However, existing E2E-AD systems typically adopt a traditional multi-task framework, addressing perception, prediction, and planning tasks through separate task-specific heads. Despite being trained in a fully differentiable manner, they still encounter issues with task coordination, and the system complexity remains high. In this work, we introduce DiffAD, a novel diffusion probabilistic model that redefines autonomous driving as a conditional image generation task. By rasterizing heterogeneous targets onto a unified bird's-eye view (BEV) and modeling their latent distribution, DiffAD unifies various driving objectives and jointly optimizes all driving tasks in a single framework, significantly reducing system complexity and harmonizing task coordination. The reverse process iteratively refines the generated BEV image, resulting in more robust and realistic driving behaviors. Closed-loop evaluations in Carla demonstrate the superiority of the proposed method, achieving a new state-of-the-art Success Rate and Driving Score.

自动驾驶扩散模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。