通过频域滤波优化扩散模型噪声,提升视频生成质量与速度
FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise
- 在频域对噪声进行精细化过滤,保持标准高斯分布特性
- 在VBench上实现最高综合评分,质量与语义表现最佳
- 引入中间步扰动采样,显著降低推理时间,适合实时应用
文本驱动的视频生成得益于扩散模型的发展。除了训练和采样阶段,近期研究关注扩散模型中的噪声先验,改进噪声先验可带来更优生成结果。一项近期方法利用傅里叶变换操作噪声,首次探索了频率域操作。然而,该方法常生成缺乏运动动态和图像细节的视频。本文对现有方法中存在的方差衰减问题进行了全面理论分析,揭示其导致细节与动态丢失的根本原因。针对噪声分布对生成质量的关键影响,我们提出FreqPrior,一种新型噪声初始化策略,通过频域滤波精细调控不同频率信号,同时维持噪声先验分布接近标准高斯分布。此外,我们设计了一种部分采样过程,在寻找噪声先验时对潜在表示于中间时间步进行扰动,显著减少推理时间且不牺牲质量。在VBench上的大量实验表明,本方法在质量与语义评估中均取得最高分,综合得分最优,充分验证了所提噪声先验的优势。
原文摘要 · Abstract (English)
Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise priors of diffusion models, as improved noise priors yield better generation results. One recent approach employs the Fourier transform to manipulate noise, marking the initial exploration of frequency operations in this context. However, it often generates videos that lack motion dynamics and imaging details. In this work, we provide a comprehensive theoretical analysis of the variance decay issue present in existing methods, contributing to the loss of details and motion dynamics. Recognizing the critical impact of noise distribution on generation quality, we introduce FreqPrior, a novel noise initialization strategy that refines noise in the frequency domain. Our method features a novel filtering technique designed to address different frequency signals while maintaining the noise prior distribution that closely approximates a standard Gaussian distribution. Additionally, we propose a partial sampling process by perturbing the latent at an intermediate timestep during finding the noise prior, significantly reducing inference time without compromising quality. Extensive experiments on VBench demonstrate that our method achieves the highest scores in both quality and semantic assessments, resulting in the best overall total score. These results highlight the superiority of our proposed noise prior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。