arXiv:2504.12532cs.LGcond-mat.dis-nn2025-04ICLR被引 26

噪声让扩散模型学会泛化,关键在训练目标的方差设计。

Generalization through variance: how noise shapes inductive biases in diffusion models

  • 用路径积分方法分析扩散模型学习到的分布,揭示其受噪声目标影响。
  • 模型实际学习的分布与训练数据相似,但填补了原始数据的空白区域。
  • 该泛化机制源于训练时噪声目标的协方差结构,适合研究生成模型原理者阅读。

扩散模型为何能泛化于训练集之外仍不明确。尽管其最优解是训练分布的得分函数,且网络表达能力足够强,但训练目标本身并非精确的得分函数,而是一个仅在期望下等于真实得分的噪声量。本文提出‘通过方差实现泛化’的理论解释:利用物理启发的路径积分方法,分析典型欠拟合与过拟合扩散模型所学分布。结果表明,模型实际学习的采样分布与训练分布相似,但在数据缺失区域被填补;这种归纳偏置源自训练中噪声目标的协方差结构,并可与特征相关的归纳偏置相互作用。

原文摘要 · Abstract (English)

How diffusion models generalize beyond their training set is not known, and is somewhat mysterious given two facts: the optimum of the denoising score matching (DSM) objective usually used to train diffusion models is the score function of the training distribution; and the networks usually used to learn the score function are expressive enough to learn this score to high accuracy. We claim that a certain feature of the DSM objective -- the fact that its target is not the training distribution's score, but a noisy quantity only equal to it in expectation -- strongly impacts whether and to what extent diffusion models generalize. In this paper, we develop a mathematical theory that partly explains this 'generalization through variance' phenomenon. Our theoretical analysis exploits a physics-inspired path integral approach to compute the distributions typically learned by a few paradigmatic under- and overparameterized diffusion models. We find that the distributions diffusion models effectively learn to sample from resemble their training distributions, but with 'gaps' filled in, and that this inductive bias is due to the covariance structure of the noisy target used during training. We also characterize how this inductive bias interacts with feature-related inductive biases.

扩散模型泛化能力生成模型归纳偏置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。