统一生成与规划的驾驶模型,用几何信息提升未来预测和决策质量。
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
- 基于多视角图像和语言指令,联合生成深度图、视频与动作
- 在Navsim v1/v2上分别达89.2 PDMS和88.7 EPDMS,优于现有方法
- 显式学习深度图增强视觉想象与规划鲁棒性,适合自动驾驶研究
近期世界-动作模型(WAM)兴起,融合视觉-语言-动作模型与世界模型的能力,实现统一推理与时空建模。但现有方法多依赖2D外观或隐表示,缺乏对物理世界中物体空间结构的几何约束。本文提出DriveDreamer-Policy,一种统一的驾驶世界-动作模型,集成深度生成、未来视频生成与运动规划于单一模块化架构。该模型利用大语言模型处理语言指令、多视角图像与动作,随后通过三个轻量生成器分别输出深度图、未来视频与动作。通过学习几何感知的世界表示,并用于引导未来预测与规划,模型生成更连贯的虚拟未来并做出更合理的驾驶决策,同时保持模块化与可控延迟。在Navsim v1和v2基准上的实验表明,该模型在闭环规划与世界生成任务中表现优异,分别取得89.2 PDMS与88.7 EPDMS的性能,超越现有基于世界模型的方法,且生成的未来视频与深度图质量更高。消融实验证明,显式深度学习对视频想象与规划鲁棒性具有互补增益。
原文摘要 · Abstract (English)
Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However, existing WAM approaches often focus on modeling 2D appearance or latent representations, with limited geometric grounding-an essential element for embodied systems operating in the physical world. We present DriveDreamer-Policy, a unified driving world-action model that integrates depth generation, future video generation, and motion planning within a single modular architecture. The model employs a large language model to process language instructions, multi-view images, and actions, followed by three lightweight generators that produce depth, future video, and actions. By learning a geometry-aware world representation and using it to guide both future prediction and planning within a unified framework, the proposed model produces more coherent imagined futures and more informed driving actions, while maintaining modularity and controllable latency. Experiments on the Navsim v1 and v2 benchmarks demonstrate that DriveDreamer-Policy achieves strong performance on both closed-loop planning and world generation tasks. In particular, our model reaches 89.2 PDMS on Navsim v1 and 88.7 EPDMS on Navsim v2, outperforming existing world-model-based approaches while producing higher-quality future video and depth predictions. Ablation studies further show that explicit depth learning provides complementary benefits to video imagination and improves planning robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。