新方法让扩散模型保留频谱结构,生成更高效多样。
Preserving Spectral Structure and Statistics in Diffusion Models
- 在频谱空间设计新前向与反向过程,不破坏结构信息。
- 生成图像质量更高,计算量降低,多样性更强。
- 适合追求高效高质量图像生成的研究者使用。
标准扩散模型将数据完全破坏为无信息的白噪声,导致反向去噪过程复杂且计算量大。本文提出一种在数学可处理的频谱空间中构建新前向与反向过程的方法。不同于基于像素的模型,该方法的前向过程收敛至一个有信息量的高斯先验 N(μ̂, Σ̂),而非白噪声。所提方法称为预保留频谱结构与统计(PreSS),在保持终态结构信号完整的同时,引导频谱成分趋向该有信息先验。这为反向过程提供了合理起点,支持基于保留频谱结构的高质量图像重建,并维持高生成多样性。在CIFAR-10、CelebA和CelebA-HQ上的实验表明,相比像素基扩散模型,该方法显著降低了计算复杂度,提升了视觉多样性,减少了漂移,使扩散过程更平滑。
原文摘要 · Abstract (English)
Standard diffusion models (DMs) rely on the total destruction of data into non-informative white noise, forcing the backward process to denoise from a fully unstructured noise state. While ensuring diversity, this results in a cumbersome and computationally intensive image generation task. We address this challenge by proposing new forward and backward process within a mathematically tractable spectral space. Unlike pixel-based DMs, our forward process converges towards an informative Gaussian prior N(mu_hat,Sigma_hat) rather than white noise. Our method, termed Preserving Spectral Structure and Statistics (PreSS) in diffusion models, guides spectral components toward this informative prior while ensuring that corresponding structural signals remain intact at terminal time. This provides a principled starting point for the backward process, enabling high-quality image reconstruction that builds upon preserved spectral structure while maintaining high generative diversity. Experimental results on CIFAR-10, CelebA and CelebA-HQ demonstrate significant reductions in computational complexity, improved visual diversity, less drift, and a smoother diffusion process compared to pixel-based DMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。