arXiv:2512.23421cs.CV2025-12被引 43

将视频生成与路径规划统一在潜空间中,提升自动驾驶预测与决策一致性。

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World

  • 用潜空间统一视频生成与路径规划,确保未来画面与轨迹同步。
  • 在FID上比最佳方法高33.3%,FVD提升1.8%,导航规划达新纪录。
  • 适合自动驾驶系统研发者,尤其关注高保真预测与可靠决策的场景。

世界模型在自动驾驶中日益重要,能学习场景随时间演变以应对真实世界的长尾挑战。然而现有方法仍将世界模型局限于有限角色:尽管架构看似统一,但视频预测与运动规划仍为分离过程。为此,我们提出DriveLaW,一种统一视频生成与运动规划的新范式。通过直接将视频生成器的潜表示注入规划器,DriveLaW确保高保真未来生成与可靠轨迹规划之间的内在一致性。具体而言,DriveLaW包含两个核心组件:驱动视频生成的DriveLaW-Video,以及基于扩散模型的DriveLaW-Act,后者从DriveLaW-Video的潜表示中生成一致可靠的轨迹。两者通过三阶段渐进训练策略联合优化。该统一范式的强大性能体现在两项任务上的新纪录:不仅在视频预测上显著超越现有最佳方法(FID提升33.3%,FVD提升1.8%),还在NAVSIM规划基准上创下新高。

原文摘要 · Abstract (English)

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate within ostensibly unified architectures that still keep world prediction and motion planning as decoupled processes. To bridge this gap, we propose DriveLaW, a novel paradigm that unifies video generation and motion planning. By directly injecting the latent representation from its video generator into the planner, DriveLaW ensures inherent consistency between high-fidelity future generation and reliable trajectory planning. Specifically, DriveLaW consists of two core components: DriveLaW-Video, our powerful world model that generates high-fidelity forecasting with expressive latent representations, and DriveLaW-Act, a diffusion planner that generates consistent and reliable trajectories from the latent of DriveLaW-Video, with both components optimized by a three-stage progressive training strategy. The power of our unified paradigm is demonstrated by new state-of-the-art results across both tasks. DriveLaW not only advances video prediction significantly, surpassing best-performing work by 33.3% in FID and 1.8% in FVD, but also achieves a new record on the NAVSIM planning benchmark.

自动驾驶视频生成潜空间规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。