arXiv:2412.05899cs.CV2024-12被引 6

四步生成高质量视频,比现有方法更快更稳。

Accelerating Video Diffusion Models via Distribution Matching

  • 用视频GAN损失与2D得分分布匹配,蒸馏出少步生成器。
  • 仅4步采样即超越现有方法,保持甚至提升画质。
  • 适合需要快速生成视频的落地场景,如内容创作。

生成模型,尤其是扩散模型,在图像、视频和3D资产等多模态数据合成中取得了显著成功。然而,当前扩散模型计算开销大,常需大量采样步骤,限制了其在视频生成等实际应用中的使用。本文提出一种新型扩散蒸馏与分布匹配框架,显著减少推理步数,同时保持甚至提升生成质量。方法通过结合视频GAN损失和新颖的2D得分分布匹配损失,将预训练扩散模型蒸馏为高效的少步生成器,特别针对视频生成任务。利用去噪GAN判别器从真实数据中蒸馏,并借助预训练图像扩散模型提升帧质量和提示遵循能力。实验以AnimateDiff为教师模型,证明该方法在仅4步采样下即可实现优于现有技术的性能。

原文摘要 · Abstract (English)

Generative models, particularly diffusion models, have made significant success in data synthesis across various modalities, including images, videos, and 3D assets. However, current diffusion models are computationally intensive, often requiring numerous sampling steps that limit their practical application, especially in video generation. This work introduces a novel framework for diffusion distillation and distribution matching that dramatically reduces the number of inference steps while maintaining-and potentially improving-generation quality. Our approach focuses on distilling pre-trained diffusion models into a more efficient few-step generator, specifically targeting video generation. By leveraging a combination of video GAN loss and a novel 2D score distribution matching loss, we demonstrate the potential to generate high-quality video frames with substantially fewer sampling steps. To be specific, the proposed method incorporates a denoising GAN discriminator to distil from the real data and a pre-trained image diffusion model to enhance the frame quality and the prompt-following capabilities. Experimental results using AnimateDiff as the teacher model showcase the method's effectiveness, achieving superior performance in just four sampling steps compared to existing techniques.

视频生成扩散模型少步采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。