arXiv:2503.03206cs.LGcs.CV2025-03NeurIPS被引 23

揭示扩散模型学习中频谱偏差的普适规律:粗结构先学,细节后成。

An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models

  • 基于高斯等价原理,推导出线性与卷积去噪器的梯度流解析解。
  • 发现学习时间与特征方差成反比,高频细节学习速度远慢于低频结构。
  • 实验证明局部卷积会改变学习动态,提示其独特归纳偏置需深入研究。

我们构建了一个解析框架,用于理解扩散模型训练过程中生成分布的演化。通过高斯等价原理,求解了线性与卷积去噪器的全批量梯度流动力学,并积分得到概率流微分方程,获得生成分布的解析表达式。理论揭示了一种普遍的逆方差频谱定律:某特征模式匹配目标方差所需时间τ与特征值λ⁻¹成正比,因此高方差(粗糙)结构比低方差(精细)细节早数个数量级被掌握。将分析扩展至深层线性网络和循环全宽卷积显示,权值共享仅乘以学习率——加速但不消除偏差;而局部卷积则引入定性不同的偏差。在高斯与自然图像数据集上的实验表明,该频谱定律在基于MLP的UNet中依然成立。然而,卷积型U-Net表现出多个模式近乎同时快速涌现,暗示局部卷积重塑了学习动态。这些结果强调数据协方差决定了扩散模型学习的顺序与速度,呼吁深入探究局部卷积带来的独特归纳偏置。

原文摘要 · Abstract (English)

We develop an analytical framework for understanding how the generated distribution evolves during diffusion model training. Leveraging a Gaussian-equivalence principle, we solve the full-batch gradient-flow dynamics of linear and convolutional denoisers and integrate the resulting probability-flow ODE, yielding analytic expressions for the generated distribution. The theory exposes a universal inverse-variance spectral law: the time for an eigen- or Fourier mode to match its target variance scales as $τ\proptoλ^{-1}$, so high-variance (coarse) structure is mastered orders of magnitude sooner than low-variance (fine) detail. Extending the analysis to deep linear networks and circulant full-width convolutions shows that weight sharing merely multiplies learning rates -- accelerating but not eliminating the bias -- whereas local convolution introduces a qualitatively different bias. Experiments on Gaussian and natural-image datasets confirm the spectral law persists in deep MLP-based UNet. Convolutional U-Nets, however, display rapid near-simultaneous emergence of many modes, implicating local convolution in reshaping learning dynamics. These results underscore how data covariance governs the order and speed with which diffusion models learn, and they call for deeper investigation of the unique inductive biases introduced by local convolution.

扩散模型频谱偏差学习动态卷积结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。