arXiv:2602.01814cs.CV2026-02被引 2

用渐进引导蒸馏,6步生成高质量视频。

GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation

  • 教师模型逐步指导学生模型增大采样步长。
  • 48步减至6步,视频质量仍达VBench基准水平。
  • 适合需要快速生成高清视频的场景。

扩散模型在视频生成中取得显著进展,但去噪过程的高计算成本仍是主要瓶颈。现有方法虽能减少扩散步骤,但在视频生成中常导致质量明显下降。本文提出引导渐进蒸馏(GPD)框架,加速扩散过程以实现快速且高质量的视频生成。GPD引入一种新型训练策略:教师模型逐步引导学生模型使用更大的采样步长。该框架包含两个关键组件:(1) 在线生成的训练目标,降低优化难度并提升计算效率;(2) 隐空间中的频域约束,有助于保留细粒度细节和时序动态。应用于Wan2.1模型,GPD将采样步数从48降至6,同时在VBench上保持竞争力的视觉质量。相较于现有蒸馏方法,GPD在流程简洁性和质量保持方面均具明显优势。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches have shown promise in reducing the number of diffusion steps, but they often suffer from significant quality degradation when applied to video generation. We propose Guided Progressive Distillation (GPD), a framework that accelerates the diffusion process for fast and high-quality video generation. GPD introduces a novel training strategy in which a teacher model progressively guides a student model to operate with larger step sizes. The framework consists of two key components: (1) an online-generated training target that reduces optimization difficulty while improving computational efficiency, and (2) frequency-domain constraints in the latent space that promote the preservation of fine-grained details and temporal dynamics. Applied to the Wan2.1 model, GPD reduces the number of sampling steps from 48 to 6 while maintaining competitive visual quality on VBench. Compared with existing distillation methods, GPD demonstrates clear advantages in both pipeline simplicity and quality preservation.

视频生成扩散模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。