arXiv:2412.04842cs.CV2024-12ICCV被引 27

统一框架生成长时多视角驾驶视频,控制更精准

UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving

  • 融合单/多视角数据训练扩散模型,增强跨帧跨视角一致性
  • 在FID和FVD上分别提升48.2%和35.2%,生成质量显著提高
  • 支持文本、图像、视频等多种输入,适用于自动驾驶仿真

构建多样化且真实的驾驶场景对提升自动驾驶系统的感知与规划能力至关重要。然而,生成长时间、多视角一致的驾驶视频仍面临巨大挑战。为此,我们提出UniMLVG——一个统一框架,可在精确控制下生成长时街景多视角视频。通过将单视角与多视角驾驶视频整合进训练数据,该方法在三个阶段中对基于DiT的扩散模型进行更新,引入跨帧与跨视角模块,并采用多目标训练策略,显著提升了生成内容的多样性和质量。尤为重要的是,我们提出一种创新的显式视角建模方法,有效改善了运动过渡的一致性。UniMLVG可处理多种输入参考格式(如文本、图像或视频),并根据3D边界框或帧级文本描述等条件约束生成高质量多视角视频。相较于具有类似能力的最佳模型,本框架在FID上提升48.2%,在FVD上提升35.2%。

原文摘要 · Abstract (English)

The creation of diverse and realistic driving scenarios has become essential to enhance perception and planning capabilities of the autonomous driving system. However, generating long-duration, surround-view consistent driving videos remains a significant challenge. To address this, we present UniMLVG, a unified framework designed to generate extended street multi-perspective videos under precise control. By integrating single- and multi-view driving videos into the training data, our approach updates a DiT-based diffusion model equipped with cross-frame and cross-view modules across three stages with multi training objectives, substantially boosting the diversity and quality of generated visual content. Importantly, we propose an innovative explicit viewpoint modeling approach for multi-view video generation to effectively improve motion transition consistency. Capable of handling various input reference formats (e.g., text, images, or video), our UniMLVG generates high-quality multi-view videos according to the corresponding condition constraints such as 3D bounding boxes or frame-level text descriptions. Compared to the best models with similar capabilities, our framework achieves improvements of 48.2% in FID and 35.2% in FVD.

自动驾驶多视角生成扩散模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。