arXiv:2505.18650cs.CV2025-05被引 7

提出可同时预测未来动作与视频的端到端驾驶世界模型

ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos

  • 用隐式动作模块和扩散模型联合学习状态与动作动态
  • 在Nuscenes上实现最佳视频一致性和动作预测准确率
  • 适合自动驾驶长时序规划与决策研究者

真实驾驶需观察环境、预判未来并做出决策,这与世界模型理解环境并预测未来的能力高度契合。然而现有自动驾驶世界模型多为显式构建,仅能通过给定等长动作序列生成视频,且忽略动作动态规律。为此,本文提出ProphetDWM,一种端到端驾驶世界模型,可联合预测未来视频与动作。该模型包含动作模块,用于从当前到未来的动作序列中学习隐式动作;以及基于扩散模型的转换模块,用于学习状态分布。模型通过联合学习:在有限状态条件下学习隐式动作,并预测动作与视频,从而连接动作动态与状态演化,实现长期未来预测。在Nuscenes数据集上的视频生成与动作预测任务中,相比现有最优方法,本方法在视频一致性与动作预测准确率上均表现最佳,同时支持高质量长时序视频与动作生成。

原文摘要 · Abstract (English)

Real-world driving requires people to observe the current environment, anticipate the future, and make appropriate driving decisions. This requirement is aligned well with the capabilities of world models, which understand the environment and predict the future. However, recent world models in autonomous driving are built explicitly, where they could predict the future by controllable driving video generation. We argue that driving world models should have two additional abilities: action control and action prediction. Following this line, previous methods are limited because they predict the video requires given actions of the same length as the video and ignore the dynamical action laws. To address these issues, we propose ProphetDWM, a novel end-to-end driving world model that jointly predicts future videos and actions. Our world model has an action module to learn latent action from the present to the future period by giving the action sequence and observations. And a diffusion-model-based transition module to learn the state distribution. The model is jointly trained by learning latent actions given finite states and predicting action and video. The joint learning connects the action dynamics and states and enables long-term future prediction. We evaluate our method in video generation and action prediction tasks on the Nuscenes dataset. Compared to the state-of-the-art methods, our method achieves the best video consistency and best action prediction accuracy, while also enabling high-quality long-term video and action generation.

自动驾驶世界模型动作预测扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。