arXiv:2410.24060cs.LGcs.CV2024-10NeurIPS被引 45

发现扩散模型泛化能力源于对数据协方差结构的隐式学习

Understanding Generalizability of Diffusion Models Requires Rethinking the Hidden Gaussian Structure

  • 通过分析得分函数的线性化,揭示扩散模型在泛化阶段趋向线性
  • 线性模型近似最优去噪器,对应训练数据的高斯分布特性
  • 小模型或训练初期即显现此偏置,解释真实扩散模型强泛化

本文研究扩散模型的泛化能力,聚焦于其学习得分函数所隐藏的性质——这些得分函数本质上是多噪声水平下训练的深度去噪器。我们观察到,当扩散模型从记忆转向泛化时,其非线性去噪器表现出越来越强的线性特征。由此,我们研究了非线性扩散模型的线性对应物,即训练以匹配非线性去噪器函数映射的一系列线性模型。令人惊讶的是,这些线性去噪器近似于由训练数据经验均值与协方差定义的多元高斯分布的最优去噪器。这表明扩散模型具有捕捉并利用训练数据协方差结构的归纳偏置。我们实证验证了这一偏置是扩散模型在泛化阶段的独特属性,尤其在模型容量相对训练集大小较小时愈发明显。当模型高度过参数化时,该偏置在训练初期便出现,早于模型完全记忆训练数据。本研究为近期真实扩散模型中观察到的强泛化现象提供了关键洞察。

原文摘要 · Abstract (English)

In this work, we study the generalizability of diffusion models by looking into the hidden properties of the learned score functions, which are essentially a series of deep denoisers trained on various noise levels. We observe that as diffusion models transition from memorization to generalization, their corresponding nonlinear diffusion denoisers exhibit increasing linearity. This discovery leads us to investigate the linear counterparts of the nonlinear diffusion models, which are a series of linear models trained to match the function mappings of the nonlinear diffusion denoisers. Surprisingly, these linear denoisers are approximately the optimal denoisers for a multivariate Gaussian distribution characterized by the empirical mean and covariance of the training dataset. This finding implies that diffusion models have the inductive bias towards capturing and utilizing the Gaussian structure (covariance information) of the training dataset for data generation. We empirically demonstrate that this inductive bias is a unique property of diffusion models in the generalization regime, which becomes increasingly evident when the model's capacity is relatively small compared to the training dataset size. In the case that the model is highly overparameterized, this inductive bias emerges during the initial training phases before the model fully memorizes its training data. Our study provides crucial insights into understanding the notable strong generalization phenomenon recently observed in real-world diffusion models.

扩散模型泛化能力协方差结构归纳偏置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。