统一建模感知与规划,提升自动驾驶端到端鲁棒性
UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving

- 将感知与规划统一在扩散模型中,通过迭代去噪联合优化
- 在多个基准上达到最优性能,超越现有判别式与扩散方法
- 适合追求高鲁棒性的自动驾驶端到端系统研究者
扩散模型在端到端自动驾驶的多模态规划中展现出强大潜力。然而,现有方法大多仅将扩散模型用于规划模块,依赖独立判别式感知网络的固定输出,导致感知误差传递至规划器,增加优化难度并降低鲁棒性。为此,我们提出UniTeD——一种统一时序扩散框架,通过共享生成空间中的迭代去噪,联合建模感知与规划。双向信息交互促进任务间的相互优化,并通过噪声条件下的多任务训练提升鲁棒性。进一步引入时序上下文,设计时序转换模块(TTM)解决历史与当前帧间噪声水平不匹配问题。同时提出锚点刷新策略(ARS),缓解稀疏扩散框架中常见的训练-推理分布偏移。无需额外组件,UniTeD在多个基准上达到领先性能,优于近期判别式端到端方法及基于扩散的规划方案。
原文摘要 · Abstract (English)
Diffusion models have shown strong potential for multi-modal planning in end-to-end autonomous driving. However, most existing methods confine diffusion to the planning module, conditioning on fixed outputs from separate discriminative perception networks. This decoupled design propagates perception errors to the planner, increasing optimization difficulty and reducing robustness. To overcome these limitations, we propose UniTeD, a Unified Temporal Diffusion framework that jointly models perception and planning through iterative denoising in a shared generative space. By enabling bidirectional information exchange, the framework facilitates mutual refinement between tasks and improves robustness via noise-conditioned multi-task training. We further extend this unified diffusion paradigm to a streaming setting by incorporating temporal context. A Temporal Transition Module (TTM) is introduced to resolve the noise-level mismatch between historical and current frames. In addition, we propose an Anchor Refresh Strategy (ARS) to alleviate the training-inference distribution shift commonly observed in sparse diffusion-based end-to-end driving frameworks. Without bells and whistles, UniTeD achieves state-of-the-art performance across multiple benchmarks, surpassing both recent discriminative end-to-end methods and diffusion-based planning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。