不训练即可加速视频生成,质量不降反升。
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
- 动态复用特征,兼顾时序连续与细节差异。
- 利用无条件与有条件特征冗余,提速1.67倍。
- 适配主流视频扩散模型,适合实时生成场景。
本文提出FasterCache,一种无需训练的视频扩散模型加速方法,可实现高质量生成。通过分析现有缓存策略发现,直接复用相邻步特征会因丢失细微变化而降低视频质量。我们首次研究了无分类器指导(CFG)的加速潜力,揭示同一时间步内条件与无条件特征存在显著冗余。基于此,提出动态特征复用策略以保持特征区分度与时间连续性,并设计CFG-Cache优化条件与无条件输出的复用,进一步提升推理速度而不损失质量。在Vchitect-2.0等最新视频扩散模型上验证,FasterCache实现最高1.67倍加速,且视频质量与基线相当,同时优于现有方法在速度与质量上的综合表现。
原文摘要 · Abstract (English)
In this paper, we present \textbf{\textit{FasterCache}}, a novel training-free strategy designed to accelerate the inference of video diffusion models with high-quality generation. By analyzing existing cache-based methods, we observe that \textit{directly reusing adjacent-step features degrades video quality due to the loss of subtle variations}. We further perform a pioneering investigation of the acceleration potential of classifier-free guidance (CFG) and reveal significant redundancy between conditional and unconditional features within the same timestep. Capitalizing on these observations, we introduce FasterCache to substantially accelerate diffusion-based video generation. Our key contributions include a dynamic feature reuse strategy that preserves both feature distinction and temporal continuity, and CFG-Cache which optimizes the reuse of conditional and unconditional outputs to further enhance inference speed without compromising video quality. We empirically evaluate FasterCache on recent video diffusion models. Experimental results show that FasterCache can significantly accelerate video generation (\eg 1.67$\times$ speedup on Vchitect-2.0) while keeping video quality comparable to the baseline, and consistently outperform existing methods in both inference speed and video quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。