统一驾驶世界模型,让车辆能同时理解场景结构、外观和动态变化。
UniDWM: Towards a Unified Driving World Model via Multifaceted Representation Learning
- 构建融合几何、纹理与动态的潜在世界表征,实现感知-预测-规划统一推理。
- 通过联合重建与条件扩散生成,实现4D场景重建与未来演化预测。
- 适用于自动驾驶系统研发者,尤其关注端到端智能决策的团队。
在复杂驾驶环境中实现可靠高效的路径规划,需要模型能够综合理解场景的几何、外观与动态特性。本文提出UniDWM,一种通过多维度表征学习实现统一驾驶世界建模的新方法。UniDWM构建了一个兼具结构与动态感知的潜在世界表示,作为物理合理的状态空间,支持感知、预测与规划的一致性推理。具体而言,联合重建路径学习恢复场景结构(包括几何与视觉纹理),而协同生成框架则利用条件扩散变换器,在潜在空间中预测未来世界演化。此外,我们证明UniDWM可视为变分自编码器(VAE)的一种变体,为多维度表征学习提供理论指导。大量实验表明,UniDWM在轨迹规划、4D重建与生成任务中均表现优异,凸显多维度世界表征在统一驾驶智能中的潜力。代码将公开于 https://github.com/Say2L/UniDWM。
原文摘要 · Abstract (English)
Achieving reliable and efficient planning in complex driving environments requires a model that can reason over the scene's geometry, appearance, and dynamics. We present UniDWM, a unified driving world model that advances autonomous driving through multifaceted representation learning. UniDWM constructs a structure- and dynamic-aware latent world representation that serves as a physically grounded state space, enabling consistent reasoning across perception, prediction, and planning. Specifically, a joint reconstruction pathway learns to recover the scene's structure, including geometry and visual texture, while a collaborative generation framework leverages a conditional diffusion transformer to forecast future world evolution within the latent space. Furthermore, we show that our UniDWM can be deemed as a variation of VAE, which provides theoretical guidance for the multifaceted representation learning. Extensive experiments demonstrate the effectiveness of UniDWM in trajectory planning, 4D reconstruction and generation, highlighting the potential of multifaceted world representations as a foundation for unified driving intelligence. The code will be publicly available at https://github.com/Say2L/UniDWM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。