用随机矩阵理论解释扩散模型为何不同数据集训练结果相似
A Random Matrix Theory Perspective on the Consistency of Diffusion Models
- 基于随机矩阵理论分析数据分布对去噪器的影响
- 发现小样本导致低方差方向过度收缩,样本趋近数据均值
- 揭示跨数据集差异的三个关键因素,适合研究生成稳定性者
在不同且不重叠的数据子集上训练的扩散模型,使用相同噪声种子时往往生成高度相似的图像。我们将其归因于一个简单的线性效应:各数据划分共享高斯统计特性,已能预测大部分生成结果。为此,我们构建了随机矩阵理论(RMT)框架,量化有限数据如何影响线性设定下学习到的去噪器与采样映射的期望和方差。对于期望,采样变异性通过自洽关系 $σ^2 o κ(σ^2)$ 实现噪声水平的重归一化,解释了为何有限数据会过度压缩低方差方向并使样本趋向数据均值。对于方差,公式揭示了跨划分分歧的三大因素:特征模态的各向异性、输入间的异质性,以及整体随数据集规模缩放的特性。进一步将确定性等价工具扩展至分数阶矩阵幂,可分析完整采样轨迹。该理论精确预测线性扩散模型行为,并在非记忆化区域的UNet与DiT架构上验证,定位了样本在不同数据划分下的偏离位置与方式。为扩散训练的可复现性提供了原理性基准,将数据谱特性与生成输出稳定性关联。
原文摘要 · Abstract (English)
Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed. We trace this consistency to a simple linear effect: the shared Gaussian statistics across splits already predict much of the generated images. To formalize this, we develop a random matrix theory (RMT) framework that quantifies how finite datasets shape the expectation and variance of the learned denoiser and sampling map in the linear setting. For expectations, sampling variability acts as a renormalization of the noise level through a self-consistent relation $σ^2 \mapsto κ(σ^2)$, explaining why limited data overshrink low-variance directions and pull samples toward the dataset mean. For fluctuations, our variance formulas reveal three key factors behind cross-split disagreement: \textit{anisotropy} across eigenmodes, \textit{inhomogeneity} across inputs, and overall scaling with dataset size. Extending deterministic-equivalence tools to fractional matrix powers further allows us to analyze entire sampling trajectories. The theory sharply predicts the behavior of linear diffusion models, and we validate its predictions on UNet and DiT architectures in their non-memorization regime, identifying where and how samples deviates across training data split. This provides a principled baseline for reproducibility in diffusion training, linking spectral properties of data to the stability of generative outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。