arXiv:2411.13807cs.CV2024-11ICCV被引 98

MagicDrive-V2实现高分辨率长时序自动驾驶视频生成与精准几何控制。

MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control

  • 采用时空条件编码与MVDiT块,支持多视角生成和精确几何控制。
  • 生成视频分辨率提升3.3倍、帧数达4倍,超越当前最先进水平。
  • 适用于需要复杂场景理解的自动驾驶仿真与测试场景。

扩散模型的快速发展显著提升了视频合成能力,尤其在可控视频生成方面,对自动驾驶等应用至关重要。尽管基于3D VAE的DiT已成为视频生成的标准框架,但在可控驾驶视频生成中仍面临几何控制难题,现有控制方法失效。为此,我们提出MagicDrive-V2,融合MVDiT模块与时空条件编码,实现多视角视频生成与精准几何控制。同时引入高效上下文描述获取方法,支持多样文本控制,并采用混合视频数据的渐进式训练策略,提升训练效率与泛化能力。实验表明,MagicDrive-V2可实现分辨率提升3.3倍、帧数增加4倍的多视角驾驶视频生成,具备丰富的上下文与几何控制能力,显著拓展自动驾驶领域的应用潜力。

原文摘要 · Abstract (English)

The rapid advancement of diffusion models has greatly improved video synthesis, especially in controllable video generation, which is vital for applications like autonomous driving. Although DiT with 3D VAE has become a standard framework for video generation, it introduces challenges in controllable driving video generation, especially for geometry control, rendering existing control methods ineffective. To address these issues, we propose MagicDrive-V2, a novel approach that integrates the MVDiT block and spatial-temporal conditional encoding to enable multi-view video generation and precise geometric control. Additionally, we introduce an efficient method for obtaining contextual descriptions for videos to support diverse textual control, along with a progressive training strategy using mixed video data to enhance training efficiency and generalizability. Consequently, MagicDrive-V2 enables multi-view driving video synthesis with $3.3\times$ resolution and $4\times$ frame count (compared to current SOTA), rich contextual control, and geometric controls. Extensive experiments demonstrate MagicDrive-V2's ability, unlocking broader applications in autonomous driving.

视频生成自动驾驶扩散模型几何控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。