arXiv:2502.06782cs.CV2025-02被引 20

基于多尺度DiT实现高效灵活的视频生成,支持动态控制与高帧率输出。

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

  • 采用多尺度分块学习提升视频生成效率与灵活性
  • 在128×128分辨率下实现30帧/秒流畅视频生成,训练效率提升40%
  • 支持运动强度直接控制,适合需要动态调节的创作场景

近期进展表明,扩散变换器(DiT)已成为生成建模的主流框架。在此基础上,Lumina-Next在高保真图像生成中表现卓越。然而,其在视频生成中的潜力尚未充分挖掘,主要受限于视频数据固有的时空复杂性。为此,我们提出Lumina-Video,该框架结合Next-DiT优势并针对视频合成设计定制化解决方案。Lumina-Video引入多尺度Next-DiT架构,联合学习多种分块方式,提升效率与灵活性;通过显式引入运动得分作为条件,实现对生成视频动态程度的直接控制。结合渐进式训练(逐级提高分辨率与帧率)及多源训练(自然与合成数据混合),Lumina-Video在高训练与推理效率下实现出色的美学质量与运动平滑性。此外,我们还提出基于Next-DiT的Lumina-V2A视频转音频模型,用于生成与视频同步的音效。代码已开源:https://www.github.com/Alpha-VLLM/Lumina-Video。

原文摘要 · Abstract (English)

Recent advancements have established Diffusion Transformers (DiTs) as a dominant framework in generative modeling. Building on this success, Lumina-Next achieves exceptional performance in the generation of photorealistic images with Next-DiT. However, its potential for video generation remains largely untapped, with significant challenges in modeling the spatiotemporal complexity inherent to video data. To address this, we introduce Lumina-Video, a framework that leverages the strengths of Next-DiT while introducing tailored solutions for video synthesis. Lumina-Video incorporates a Multi-scale Next-DiT architecture, which jointly learns multiple patchifications to enhance both efficiency and flexibility. By incorporating the motion score as an explicit condition, Lumina-Video also enables direct control of generated videos' dynamic degree. Combined with a progressive training scheme with increasingly higher resolution and FPS, and a multi-source training scheme with mixed natural and synthetic data, Lumina-Video achieves remarkable aesthetic quality and motion smoothness at high training and inference efficiency. We additionally propose Lumina-V2A, a video-to-audio model based on Next-DiT, to create synchronized sounds for generated videos. Codes are released at https://www.github.com/Alpha-VLLM/Lumina-Video.

视频生成扩散模型多尺度动态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。