用尺度不变噪声替代白噪声,提升图像生成速度与细节质量
Cloud Diffusion Part 1: Theory and Motivation
- 用幂律分布的尺度不变噪声替代传统白噪声
- 可加速推理过程,增强高频细节表现
- 适合追求高效高质图像生成的研究者
图像生成的扩散模型通过逐步添加白噪声并训练模型从噪声中还原信号来工作。白噪声在每个点上独立服从正态分布,均值和方差不随尺度变化。然而,自然图像的低阶统计特性表现出幂律缩放的尺度不变性,使其更接近一种强调大尺度相关性、弱化小尺度相关性的概率分布。将这种尺度不变噪声引入扩散模型,形成所谓「云扩散模型」(Cloud Diffusion Model)。我们论证此类模型可实现更快推理、更好高频细节表现以及更强可控性。后续论文将构建并训练基于尺度不变性的云扩散模型,并与经典白噪声扩散模型进行对比。
原文摘要 · Abstract (English)
Diffusion models for image generation function by progressively adding noise to an image set and training a model to separate out the signal from the noise. The noise profile used by these models is white noise -- that is, noise based on independent normal distributions at each point whose mean and variance is independent of the scale. By contrast, most natural image sets exhibit a type of scale invariance in their low-order statistical properties characterized by a power-law scaling. Consequently, natural images are closer (in a quantifiable sense) to a different probability distribution that emphasizes large scale correlations and de-emphasizes small scale correlations. These scale invariant noise profiles can be incorporated into diffusion models in place of white noise to form what we will call a ``Cloud Diffusion Model". We argue that these models can lead to faster inference, improved high-frequency details, and greater controllability. In a follow-up paper, we will build and train a Cloud Diffusion Model that uses scale invariance at a fundamental level and compare it to classic, white noise diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。