揭示扩散模型在有限数据下的泛化机制,发现数据层级结构可显著提升学习效率。
Generalization Dynamics of Linear Diffusion Models
- 基于线性神经网络与数据协方差谱分析,构建扩散模型泛化理论框架。
- 当样本数 $N < d$ 时,训练与测试损失差距大;当 $N > d$ 时,泛化误差随 $d/N$ 线性下降。
- 层级数据结构与正则化能有效防过拟合,适合研究生成模型泛化行为的学者。
扩散模型是强大的生成模型,能从复杂数据中生成高质量样本。尽管其无限数据下的行为已清晰,但有限数据下的泛化机制仍不明确。经典学习理论预测泛化所需样本复杂度随维度呈指数增长,远超实际需求。本文通过数据协方差谱视角分析扩散模型,发现真实数据的协方差常呈幂律衰减,反映其层次结构。我们建立基于线性神经网络的理论框架,在高斯假设下量化数据方差层次结构与正则化对泛化的影响。当 $N < d$ 时,训练数据未覆盖所有变化方向,导致训练与测试损失差距大;此时强层级结构、正则化和早停可缓解过拟合。当 $N > d$ 时,线性扩散模型的采样分布以 $d/N$ 的速率线性逼近最优(以 KL 散度衡量),与数据分布具体形式无关。本工作阐明了样本复杂度如何决定扩散生成模型的泛化能力。
原文摘要 · Abstract (English)
Diffusion models are powerful generative models that produce high-quality samples from complex data. While their infinite-data behavior is well understood, their generalization with finite data remains less clear. Classical learning theory predicts that generalization occurs at a sample complexity that is exponential in the dimension, far exceeding practical needs. We address this gap by analyzing diffusion models through the lens of data covariance spectra, which often follow power-law decays, reflecting the hierarchical structure of real data. To understand whether such a hierarchical structure can benefit learning in diffusion models, we develop a theoretical framework based on linear neural networks, congruent with a Gaussian hypothesis on the data. We quantify how the hierarchical organization of variance in the data and regularization impacts generalization. We find two regimes: When $N <d$, not all directions of variation are present in the training data, which results in a large gap between training and test loss. In this regime, we demonstrate how a strongly hierarchical data structure, as well as regularization and early stopping help to prevent overfitting. For $N > d$, we find that the sampling distributions of linear diffusion models approach their optimum (measured by the Kullback-Leibler divergence) linearly with $d/N$, independent of the specifics of the data distribution. Our work clarifies how sample complexity governs generalization in a simple model of diffusion-based generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。