仅用一步生成高质量视频,速度与效果双突破。
OSV: One Step is Enough for High-Quality Image to Video Generation
- 两阶段训练融合一致性蒸馏与GAN,提升效率与稳定性。
- 1步生成视频FVD达171.15,超越8步方法表现。
- 新判别器无需解码视频隐变量,适合快速生成场景。
视频扩散模型在生成高质量视频方面展现出巨大潜力,但其固有的迭代特性导致计算和时间成本高昂。尽管已有研究通过一致性蒸馏等方法减少推理步数,或采用GAN训练加速,但这些方法常在性能或训练稳定性上表现不足。本文提出一种两阶段训练框架,有效结合一致性蒸馏与GAN训练以应对上述挑战。此外,我们设计了一种新型视频判别器,无需解码视频隐变量,显著提升生成质量。所提模型仅需一步即可生成高质量视频,并可灵活扩展至多步精炼以进一步提升性能。在OpenWebVid-1M基准上的定量评估表明,我们的模型显著优于现有方法:1步生成的FVD为171.15,超过基于一致性蒸馏的方法AnimateLCM(8步,FVD 184.79),接近先进模型Stable Video Diffusion(25步,FVD 156.94)的表现。
原文摘要 · Abstract (English)
Video diffusion models have shown great potential in generating high-quality videos, making them an increasingly popular focus. However, their inherent iterative nature leads to substantial computational and time costs. While efforts have been made to accelerate video diffusion by reducing inference steps (through techniques like consistency distillation) and GAN training (these approaches often fall short in either performance or training stability). In this work, we introduce a two-stage training framework that effectively combines consistency distillation with GAN training to address these challenges. Additionally, we propose a novel video discriminator design, which eliminates the need for decoding the video latents and improves the final performance. Our model is capable of producing high-quality videos in merely one-step, with the flexibility to perform multi-step refinement for further performance enhancement. Our quantitative evaluation on the OpenWebVid-1M benchmark shows that our model significantly outperforms existing methods. Notably, our 1-step performance(FVD 171.15) exceeds the 8-step performance of the consistency distillation based method, AnimateLCM (FVD 184.79), and approaches the 25-step performance of advanced Stable Video Diffusion (FVD 156.94).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。