arXiv:2508.14327cs.CV2025-08

统一框架生成多模态多视角驾驶视频,提升自动驾驶场景建模能力。

MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer

  • 构建共享与专有组件结合的扩散Transformer,统一生成多模态数据。
  • 在真实自动驾驶数据集上实现高质量、可控制的多视角视频生成。
  • 适合自动驾驶仿真、感知系统训练等场景使用。

基于视频生成模型的城市场景合成近期在自动驾驶领域展现出巨大潜力。现有方法主要聚焦于RGB视频生成,缺乏对多模态数据(如深度图、语义图)的支持。尽管可通过多个模型分别生成不同模态,但会增加部署难度,且无法利用模态间的互补信息。为此,本文提出一种面向自动驾驶的新型多模态多视角视频生成方法。我们构建了一个由模态共享与模态专属组件组成的统一扩散Transformer模型,并通过多样化条件输入编码可控的场景结构与内容线索。该方法可在统一框架内生成多模态、多视角的驾驶场景视频。在真实自动驾驶数据集上的充分实验表明,相比现有最优方法,本方法在视频生成质量与可控性方面表现优异,同时支持多模态多视角数据生成。

原文摘要 · Abstract (English)

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack the ability to support multi-modal video generation. However, multi-modal data, such as depth maps and semantic maps, are crucial for holistic urban scene understanding in autonomous driving. Although it is feasible to use multiple models to generate different modalities, this increases the difficulty of model deployment and does not leverage complementary cues for multi-modal data generation. To address this problem, in this work, we propose a novel multi-modal multi-view video generation approach to autonomous driving. Specifically, we construct a unified diffusion transformer model composed of modal-shared components and modal-specific components. Then, we leverage diverse conditioning inputs to encode controllable scene structure and content cues into the multi-modal multi-view unified diffusion model. In this way, our approach is capable of generating multi-modal multi-view driving scene videos in a unified framework. Our thorough experiments on real-world autonomous driving dataset show that our approach achieves compelling video generation quality and controllability compared with state-of-the-art methods, while supporting multi-modal multi-view data generation.

视频生成扩散模型自动驾驶多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。