用自然图像的1/f噪声加速扩散模型采样,效果更好且更快。
Flicker-DDPM: Accelerating Denoising Diffusion via 1/f Colored Noise Injection

- 引入1/f色噪声替代传统白噪声,更贴合自然图像频谱特性。
- 在CIFAR-10上仅需3.33倍少的采样步数即达同等生成质量。
- 理论证明频谱匹配可线性化反向轨迹,解释加速机制。
我们提出一种新型扩散模型Flicker-DDPM,其在前向过程引入受自组织临界性启发的1/f(闪烁)噪声。与传统去噪扩散概率模型(DDPM)使用各向同性白噪声不同,Flicker-DDPM采用具有幂律谱的色噪声,更匹配自然图像的功率谱特性(通常满足P(k) ∝ 1/k^α)。为此,我们设计基于空间相关核σ(d) = (d + 1)^{-η}的色噪声模块,并理论上证明调节η可控制生成噪声的谱指数α,实现对不同数据集谱特性的适配。在CIFAR-10上,Flicker-DDPM以3.33倍更少的采样步数达到或超越标准DDPM基线的生成质量,每步计算开销几乎无增加。我们进一步构建频域线性理论,证明谱匹配的色噪声可线性化反向轨迹,为观察到的采样加速提供理论解释。
原文摘要 · Abstract (English)
We propose a novel diffusion model, Flicker-DDPM, which incorporates flicker (1/f) noise inspired by self-organized criticality (SOC), a widely observed phenomenon in natural systems. Unlike denoising diffusion probabilistic models (DDPMs), which employ isotropic white noise in the forward process, Flicker-DDPM adopts colored noise with power-law spectra to better match the spectral statistics of natural images, whose power spectra typically follow P(k) proportional to 1/k^α. To this end, we develop a colored-noise module based on a spatial correlation kernel, σ(d) = (d + 1)^{-η}, and theoretically establish that adjusting η controls the spectral exponent α of the generated 1/fα noise, enabling adaptation to datasets with diverse spectral characteristics. On CIFAR-10, Flicker DDPM matches or surpasses the generation quality of a standard DDPM baseline using 3.33 times fewer sampling steps, with negligible additional computational cost per step. We further develop a frequency-domain linear theory demonstrating that spectrally matched colored noise linearizes the reverse trajectory, theoretically explaining the observed sampling acceleration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。