arXiv:2504.19614cs.CV2025-04被引 12

DiVE用扩散Transformer生成高质量多视角驾驶视频,提升感知模型训练效果。

DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer

  • 基于扩散Transformer,用统一注意力与草图网络控制多模态输入
  • 在nuScenes上实现当前最佳多视角视频生成,时序与跨视角一致性高
  • 提出轻量级加速策略,生成速度提升2.62倍,适合自动驾驶数据增强

为提升3D视觉感知任务性能,收集多视角驾驶场景视频面临巨大挑战且成本高昂,生成模型成为有吸引力的替代方案。然而,现有方法生成的视频质量差、时空不一致,难以用于驱动场景下的感知任务。为此,我们提出DiVE,一种基于扩散Transformer的生成框架,可生成高保真、时序连贯、跨视角一致的多视角视频,并与鸟瞰图布局和文本描述对齐。DiVE采用统一交叉注意力机制与SketchFormer,实现对多模态数据的精准控制;引入无额外参数的视图膨胀注意力机制,保障多视角一致性。然而,在多模态约束下生成高分辨率视频仍面临双重挑战:如何在复杂多条件输入下优化分类器无关引导(CFG)配置,以及如何缓解高分辨率渲染带来的计算延迟——这两点在以往研究中均未充分探索。为此,我们提出两项创新:多控制辅助分支蒸馏(Multi-Control Auxiliary Branch Distillation),在避免高计算开销的前提下简化多条件CFG选择;分辨率渐进采样(Resolution Progressive Sampling),一种无需训练的加速策略,通过分阶段提升分辨率降低高分辨率导致的延迟。二者协同实现2.62倍加速,质量损失极小。在nuScenes数据集上的评估表明,DiVE达到多视角视频生成的最先进水平,生成结果高度逼真,具有出色的时序与跨视角一致性。

原文摘要 · Abstract (English)

Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative. Yet, the videos generated by recent works suffer from poor quality and spatiotemporal consistency, undermining their utility in advancing perception tasks under driving scenarios. To address this gap, we propose DiVE, a diffusion transformer-based generative framework meticulously engineered to produce high-fidelity, temporally coherent, and cross-view consistent multi-view videos, aligning seamlessly with bird's-eye view layouts and textual descriptions. DiVE leverages a unified cross-attention and a SketchFormer to exert precise control over multimodal data, while incorporating a view-inflated attention mechanism that adds no extra parameters, thereby guaranteeing consistency across views. Despite these advancements, synthesizing high-resolution videos under multimodal constraints introduces dual challenges: investigating the optimal classifier-free guidance coniguration under intricate multi-condition inputs and mitigating excessive computational latency in high-resolution rendering--both of which remain underexplored in prior researches. To resolve these limitations, we introduce two innovations: Multi-Control Auxiliary Branch Distillation, which streamlines multi-condition CFG selection while circumventing high computational overhead, and Resolution Progressive Sampling, a training-free acceleration strategy that staggers resolution scaling to reduce high latency due to high resolution. These innovations collectively achieve a 2.62x speedup with minimal quality degradation. Evaluated on the nuScenes dataset, DiVE achieves SOTA performance in multi-view video generation, yielding photorealistic outputs with exceptional temporal and cross-view coherence.

视频生成扩散模型自动驾驶多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。