提出高效算法实现扩散模型视频生成的无限分辨率噪声一致性
Infinite-Resolution Integral Noise Warping for Diffusion Models
- 通过布朗桥增量聚合,实现无损时间一致性
- 计算成本降低数量级,精度达无限分辨率极限
- 适用于真实视频生成,可扩展至三维场景
将预训练图像扩散模型适配为生成时序一致视频已成为重要研究方向。无需训练的噪声空间操作已被证明有效,但关键挑战在于保持高斯白噪声分布的同时引入时间一致性。Chang 等(2024)提出了基于积分噪声表示的方法,具备分布保持性,并设计了上采样算法实现该表示。然而其算法计算开销高昂。本文分析其在上采样分辨率趋于无穷时的极限行为,提出一种新算法:通过聚合多个布朗桥增量,在达到无限分辨率精度的同时,将计算成本降低数个数量级。我们理论证明并实验验证了该方法的有效性,展示了其在真实场景中的应用能力。此外,该方法可自然扩展至三维空间。
原文摘要 · Abstract (English)
Adapting pretrained image-based diffusion models to generate temporally consistent videos has become an impactful generative modeling research direction. Training-free noise-space manipulation has proven to be an effective technique, where the challenge is to preserve the Gaussian white noise distribution while adding in temporal consistency. Recently, Chang et al. (2024) formulated this problem using an integral noise representation with distribution-preserving guarantees, and proposed an upsampling-based algorithm to compute it. However, while their mathematical formulation is advantageous, the algorithm incurs a high computational cost. Through analyzing the limiting-case behavior of their algorithm as the upsampling resolution goes to infinity, we develop an alternative algorithm that, by gathering increments of multiple Brownian bridges, achieves their infinite-resolution accuracy while simultaneously reducing the computational cost by orders of magnitude. We prove and experimentally validate our theoretical claims, and demonstrate our method's effectiveness in real-world applications. We further show that our method readily extends to the 3-dimensional space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。