arXiv:2512.04459cs.CV2025-12被引 14

用扩散模型提升自动驾驶的推理与规划一致性

dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning

  • 采用双向注意力的扩散模型实现可控制的推理与规划统一
  • 在nuScenes和WOD-E2E上提升9%轨迹一致性,长尾场景RFS增6%
  • 适合关注端到端自动驾驶可靠性与可控性的研究者

自动驾驶领域日益关注分布外(OOD)驾驶场景的挑战。主流研究通过整合视觉语言模型(VLMs)来增强端到端(E2E)驾驶系统,利用其丰富的世界知识与推理能力以提升跨环境泛化性能。然而,现有大多数VLM或视觉语言代理(VLA)基于自回归(AR)模型,受限于因果注意力与顺序生成,难以保持高层推理与底层规划之间的一致性与可控性。相反,具备双向注意力的离散扩散VLM通过迭代去噪展现出更优的可控性与可靠性。基于此,我们提出dVLM-AD,一种基于扩散的视觉语言模型,统一感知、结构化推理与低层规划,用于端到端驾驶。在nuScenes和WOD-E2E上评估显示,dVLM-AD生成更一致的推理-动作对,在骨干网络较简的情况下,规划性能媲美现有驾驶VLM/VLA系统,相比AR基线在行为-轨迹一致性上提升9%,在长尾WOD-E2E场景中RFS提高6%。这些结果表明,可控且可靠的路径为可扩展端到端驾驶提供了新方向。

原文摘要 · Abstract (English)

The autonomous driving community is increasingly focused on addressing the challenges posed by out-of-distribution (OOD) driving scenarios. A dominant research trend seeks to enhance end-to-end (E2E) driving systems by integrating vision-language models (VLMs), leveraging their rich world knowledge and reasoning abilities to improve generalization across diverse environments. However, most existing VLMs or vision-language agents (VLAs) are built upon autoregressive (AR) models. In this paper, we observe that existing AR-based VLMs -- limited by causal attention and sequential token generation -- often fail to maintain consistency and controllability between high-level reasoning and low-level planning. In contrast, recent discrete diffusion VLMs equipped with bidirectional attention exhibit superior controllability and reliability through iterative denoising. Building on these observations, we introduce dVLM-AD, a diffusion-based vision-language model that unifies perception, structured reasoning, and low-level planning for end-to-end driving. Evaluated on nuScenes and WOD-E2E, dVLM-AD yields more consistent reasoning-action pairs and achieves planning performance comparable to existing driving VLM/VLA systems despite a modest backbone, outperforming AR-based baselines with a 9 percent improvement in behavior-trajectory consistency and a 6 percent increase in RFS on long-tail WOD-E2E scenarios. These results suggest a controllable and reliable pathway for scalable end-to-end driving.

自动驾驶扩散模型多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。