arXiv:2502.05179cs.CV2025-02AAAI被引 37

分两阶段生成高清视频,速度更快质量更高

FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation

论文配图:FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
图 1 · 摘自论文原文
  • 先低分辨率生成,用大模型和多步推断保证提示契合度
  • 再通过流匹配实现高低分辨率间近乎直线的轨迹过渡,细节自然且仅需少量推断步数
  • 支持中途预览调参,显著降低算力消耗与等待时间,适合实际应用

DiT模型在文本到视频生成中取得显著进展,凭借其模型容量与数据规模的可扩展性。然而,高内容与运动保真度通常需要大量参数和较多函数求值次数(NFEs)。真实且视觉吸引人的细节往往体现在高分辨率输出中,进一步加剧计算负担,尤其对单阶段DiT模型而言。为此,我们提出一种新颖的两阶段框架FlashVideo,通过在不同阶段策略性分配模型容量与NFEs,以平衡生成保真度与质量。第一阶段优先保证提示保真度,通过低分辨率生成过程使用大参数量和充足的NFEs,提升计算效率;第二阶段通过流匹配在低与高分辨率间实现近乎直线的ODE轨迹,有效生成精细细节并修复伪影,仅需极少的NFEs。为确保推理时两个独立训练阶段间的无缝衔接,我们在第二阶段训练中精心设计退化策略。定量与视觉结果表明,FlashVideo在高分辨率视频生成上达到当前最佳水平,同时具备卓越的计算效率。此外,两阶段设计使用户可在全分辨率生成前预览初始输出并调整提示,显著降低计算成本与等待时间,提升商业可行性。

原文摘要 · Abstract (English)

DiT models have achieved great success in text-to-video generation, leveraging their scalability in model capacity and data scale. High content and motion fidelity aligned with text prompts, however, often require large model parameters and a substantial number of function evaluations (NFEs). Realistic and visually appealing details are typically reflected in high-resolution outputs, further amplifying computational demands-especially for single-stage DiT models. To address these challenges, we propose a novel two-stage framework, FlashVideo, which strategically allocates model capacity and NFEs across stages to balance generation fidelity and quality. In the first stage, prompt fidelity is prioritized through a low-resolution generation process utilizing large parameters and sufficient NFEs to enhance computational efficiency. The second stage achieves a nearly straight ODE trajectory between low and high resolutions via flow matching, effectively generating fine details and fixing artifacts with minimal NFEs. To ensure a seamless connection between the two independently trained stages during inference, we carefully design degradation strategies during the second-stage training. Quantitative and visual results demonstrate that FlashVideo achieves state-of-the-art high-resolution video generation with superior computational efficiency. Additionally, the two-stage design enables users to preview the initial output and accordingly adjust the prompt before committing to full-resolution generation, thereby significantly reducing computational costs and wait times as well as enhancing commercial viability.

视频生成扩散模型高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。