arXiv:2503.13587cs.CV2025-03中稿 · ICRA被引 19

统一建模驾驶场景的4D动态演化,兼顾视觉与几何一致性。

UniFuture: A 4D Driving World Model for Future Generation and Perception

  • 将图像与深度图视为同一4D现实的耦合投影,联合建模。
  • 在nuScenes和Waymo上优于专用模型,生成高质量4D序列。
  • 适合自动驾驶中需要时空一致感知与生成的研究者。

我们提出UniFuture,一种统一的4D驾驶世界模型,用于模拟3D物理世界的动态演化。与仅关注2D像素级视频生成(缺乏几何)或静态感知(缺乏时间动态)的现有模型不同,我们的方法融合外观与几何,构建完整的4D表征。具体而言,将未来RGB图像与深度图视为同一4D现实的耦合投影,并在单一框架内联合建模。为此,我们引入双隐空间共享(DLS)机制,将视觉与几何模态映射到共享时空隐空间,隐式关联纹理与结构。此外,提出多尺度隐空间交互(MLI)机制,实现双向一致性:几何约束视觉合成以防止结构幻觉,视觉语义优化几何估计。推理时,UniFuture可从单帧当前图像生成高保真、几何一致的4D场景序列(图像-深度对)。在nuScenes和Waymo数据集上的大量实验表明,该方法在未来的生成与几何感知上均优于专用模型,验证了统一4D建模在自动驾驶中的有效性。代码已开源:https://github.com/dk-liang/UniFuture。

原文摘要 · Abstract (English)

We present UniFuture, a unified 4D Driving World Model designed to simulate the dynamic evolution of the 3D physical world. Unlike existing driving world models that focus solely on 2D pixel-level video generation (lacking geometry) or static perception (lacking temporal dynamics), our approach bridges appearance and geometry to construct a holistic 4D representation. Specifically, we treat future RGB images and depth maps as coupled projections of the same 4D reality and model them jointly within a single framework. To achieve this, we introduce a Dual-Latent Sharing (DLS) scheme, which maps visual and geometric modalities into a shared spatio-temporal latent space, implicitly entangling texture with structure. Furthermore, we propose a Multi-scale Latent Interaction (MLI) mechanism, which enforces bidirectional consistency: geometry constrains visual synthesis to prevent structural hallucinations, while visual semantics refine geometric estimation. During inference, UniFuture can forecast high-fidelity, geometrically consistent 4D scene sequences (image-depth pairs) from a single current frame. Extensive experiments on the nuScenes and Waymo datasets demonstrate that our method outperforms specialized models in both future generation and geometry perception, highlighting the efficacy of unified 4D modeling for autonomous driving. The code is available at https://github.com/dk-liang/UniFuture.

4D建模自动驾驶生成模型多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。