arXiv:2607.26004cs.CVcs.LG2026-07

用并行解码蒸馏加速图像视频生成,提升速度与多样性。

Parallel Decoding Distillation for Fast Image and Video Generation

  • 通过预测多步去噪实现每轮网络评估加速生成
  • 4-8次函数求值下达到当前最优性能
  • 适合需要快速生成且注重视频多样性的研究者

视频扩散或流模型的生成过程因缓慢的迭代采样而计算成本高昂。现有最先进加速方法严重依赖变分分数蒸馏(VSD)和对抗损失来将扩散模型压缩为少步生成器,尽管生成质量高,但训练损失难以优化,易出现模式崩溃,导致视频多样性下降、运动缺失。本文提出并行解码蒸馏(PDD),一种简化且可扩展的基于轨迹的蒸馏方法,用于加速扩散与流匹配模型的推理。其架构和训练流程兼容任意预训练模型,支持可变函数求值次数(NFE)。PDD通过每轮网络评估预测多个去噪步骤实现加速。概念上,它学习均值速度表示,而非使用JVP或有限差分近似回归其导数。在LTX-2.3 Text-to-Video/Audio、Wan 14B Text-to-Video和Qwen-Image Text-to-Image上,仅需4-8 NFE即达到最先进性能,同时显著提升生成视频多样性。

原文摘要 · Abstract (English)

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion. In this paper, we introduce Parallel Decoding Distillation (PDD), a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models. Our architecture and training procedure are compatible with any pre-trained model and support sampling with a varying number of function evaluations (NFE). PDD accelerates generation by predicting multiple denoising steps per network evaluation. Conceptually, it learns a representation of the mean velocity without regressing its derivative using JVPs or finite-difference approximations. Our method achieves SOTA performance with 4-8 NFE on LTX-2.3 Text-to-Video/Audio, Wan 14B Text-to-Video, and Qwen-Image Text-to-Image. Moreover, PDD presents a significant improvement in generated video diversity.

视频生成扩散模型加速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。