arXiv:2602.02002cs.CV2026-02被引 7

统一生成自动驾驶多模态未来观测,无需中间步骤。

UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving

  • 单阶段联合建模视频与激光雷达数据,直接生成未来多模态输出。
  • 在Waymo Open Dataset上,视频和激光雷达生成质量均超越现有方法。
  • 适合需要高精度仿真数据的自动驾驶研发团队使用。

世界模型在自动驾驶数据合成中展现出巨大潜力。然而,现有方法主要聚焦于单模态生成,通常仅处理多相机视频或激光雷达序列。本文提出UniDriveDreamer,一种面向自动驾驶的单阶段统一多模态世界模型,可直接生成多模态未来观测,无需依赖中间表示或级联模块。框架引入专用于激光雷达序列的变分自编码器(VAE)和用于多相机图像的视频VAE。为确保跨模态兼容性与训练稳定性,提出统一潜在锚定(ULA),显式对齐两种模态的潜在分布。对齐特征经融合后由扩散变换器处理,联合建模几何对应关系与时间演化。此外,结构化场景布局信息作为条件信号分别投影至各模态以引导合成。大量实验表明,UniDriveDreamer在视频与激光雷达生成方面均优于当前最优方法,并在下游任务中带来可测量的性能提升。

原文摘要 · Abstract (English)

World models have demonstrated significant promise for data synthesis in autonomous driving. However, existing methods predominantly concentrate on single-modality generation, typically focusing on either multi-camera video or LiDAR sequence synthesis. In this paper, we propose UniDriveDreamer, a single-stage unified multimodal world model for autonomous driving, which directly generates multimodal future observations without relying on intermediate representations or cascaded modules. Our framework introduces a LiDAR-specific variational autoencoder (VAE) designed to encode input LiDAR sequences, alongside a video VAE for multi-camera images. To ensure cross-modal compatibility and training stability, we propose Unified Latent Anchoring (ULA), which explicitly aligns the latent distributions of the two modalities. The aligned features are fused and processed by a diffusion transformer that jointly models their geometric correspondence and temporal evolution. Additionally, structured scene layout information is projected per modality as a conditioning signal to guide the synthesis. Extensive experiments demonstrate that UniDriveDreamer outperforms previous state-of-the-art methods in both video and LiDAR generation, while also yielding measurable improvements in downstream

自动驾驶多模态生成扩散模型世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。