统一生成自动驾驶多模态未来观测,无需中间步骤。
UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving
- 单阶段联合建模视频与激光雷达数据,直接生成未来多模态输出。
- 在Waymo Open Dataset上,视频和激光雷达生成质量均超越现有方法。
- 适合需要高精度仿真数据的自动驾驶研发团队使用。
世界模型在自动驾驶数据合成中展现出巨大潜力。然而,现有方法主要聚焦于单模态生成,通常仅处理多相机视频或激光雷达序列。本文提出UniDriveDreamer,一种面向自动驾驶的单阶段统一多模态世界模型,可直接生成多模态未来观测,无需依赖中间表示或级联模块。框架引入专用于激光雷达序列的变分自编码器(VAE)和用于多相机图像的视频VAE。为确保跨模态兼容性与训练稳定性,提出统一潜在锚定(ULA),显式对齐两种模态的潜在分布。对齐特征经融合后由扩散变换器处理,联合建模几何对应关系与时间演化。此外,结构化场景布局信息作为条件信号分别投影至各模态以引导合成。大量实验表明,UniDriveDreamer在视频与激光雷达生成方面均优于当前最优方法,并在下游任务中带来可测量的性能提升。
原文摘要 · Abstract (English)
World models have demonstrated significant promise for data synthesis in autonomous driving. However, existing methods predominantly concentrate on single-modality generation, typically focusing on either multi-camera video or LiDAR sequence synthesis. In this paper, we propose UniDriveDreamer, a single-stage unified multimodal world model for autonomous driving, which directly generates multimodal future observations without relying on intermediate representations or cascaded modules. Our framework introduces a LiDAR-specific variational autoencoder (VAE) designed to encode input LiDAR sequences, alongside a video VAE for multi-camera images. To ensure cross-modal compatibility and training stability, we propose Unified Latent Anchoring (ULA), which explicitly aligns the latent distributions of the two modalities. The aligned features are fused and processed by a diffusion transformer that jointly models their geometric correspondence and temporal evolution. Additionally, structured scene layout information is projected per modality as a conditioning signal to guide the synthesis. Extensive experiments demonstrate that UniDriveDreamer outperforms previous state-of-the-art methods in both video and LiDAR generation, while also yielding measurable improvements in downstream
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。