用视觉语言模型指导扩散模型,实现高效多模态自动驾驶决策
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
- 引入视觉语言模型增强扩散策略,实现稀疏-稠密混合表示
- 在2025自动驾驶大赛中达到45.0 PDMS,优于现有方法
- 适合关注端到端自动驾驶与多模态行为生成的研究者
端到端自动驾驶研究因可微分设计整合感知、预测与规划而备受关注,支持以最终目标为导向的联合优化。然而现有方法仍面临鸟瞰图(BEV)计算成本高、动作多样性不足以及复杂真实场景下决策不佳等问题。为此,我们提出一种新型混合稀疏-稠密扩散策略,依托视觉语言模型(VLM),称为Diff-VLA。通过稀疏扩散表示实现高效的多模态驾驶行为建模,并重新思考VLM在驾驶决策中的作用,通过智能体、地图实例与VLM输出间的深度交互提升轨迹生成引导能力。该方法在包含挑战性真实与动态合成场景的2025自动驾驶大赛中表现优异,取得45.0 PDMS得分。
原文摘要 · Abstract (English)
Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called Diff-VLA. We explore the sparse diffusion representation for efficient multi-modal driving behavior. Moreover, we rethink the effectiveness of VLM driving decision and improve the trajectory generation guidance through deep interaction across agent, map instances and VLM output. Our method shows superior performance in Autonomous Grand Challenge 2025 which contains challenging real and reactive synthetic scenarios. Our methods achieves 45.0 PDMS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。