用小波分析扩散模型的得分函数,揭示不同架构为何生成效果不同。
Where the Score Lives: A Wavelet View of Diffusion

- 基于二维正交小波基解析参数化得分函数,可解释其与数据分布的关系。
- 发现数据分布的矩信息对去噪过程影响最大,决定生成质量关键因素。
- 无需依赖具体网络结构,适用于理解CNN、U-Net等架构的生成差异。
基于得分的生成模型在过去十年中在生成多样且视觉逼真的图像方面取得了显著成功。尽管已采用多种架构(如CNN、U-Net和Transformer)作为得分近似网络,但目前对这些架构选择如何影响生成行为仍知之甚少。本文提出一种基于二维正交小波基的可解析参数化得分函数,推导出以数据分布矩表示的可解释最优得分函数。通过该参数化方法,实现无需特定架构的、基于矩的分析,揭示哪些数据分布属性最影响去噪过程。该得分模型足够灵活,能部分模拟U-Net和CNN的归纳偏置,为理解不同得分架构产生不同生成行为提供了新视角。由于得分函数可由数据分布矩表达,我们得以初步理解数据分布与得分网络如何共同作用,产生扩散模型中的观测行为。
原文摘要 · Abstract (English)
Score-based generative models have had remarkable success over the last decade in generating a diverse set of visually plausible images. A variety of architectures including CNNs, U-Nets, and Transformers have been used as the score-approximation network in such diffusion modeling; however, to date, relatively little is known about how these architectural choices impact generative behavior. In this work, to provide insight into this area, we propose an analytically solvable parameterization of the score function using an expansion in a 2D orthogonal wavelet basis. In particular, we derive interpretable optimal score functions in terms of the moments of the data distribution. We use this parametrization to provide an architecture-agnostic, moment-based analysis that reveals which attributes of the data distribution tend to matter most for denoising. Our score machine is flexible enough to partially mimic the relevant inductive biases of multiple architectures, including U-Nets, and CNNs, taking a step towards understanding why different score architectures can exhibit distinct generative behavior. Since our score is solvable in terms of the moments of the data, we can begin to understand how the data distribution interacts with the score network to produce the behavior we observe in diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。