揭示扩散模型从简单到复杂的统计学习规律。
A theory of learning data statistics in diffusion models, from easy to hard
- 通过最小数据模型验证模型先学成对统计,再学高阶相关。
- 发现扩散信息指数决定学习复杂度:二阶为线性,四阶需立方样本量。
- 当高阶与低阶共享潜在结构时,四阶学习可降为线性,适合研究生成机制者。
尽管扩散模型已成为强大的生成模型,但其学习动态仍不清晰。我们首先通过实证发现,标准扩散模型在自然图像上训练时表现出分布简化偏倚,优先学习简单的成对输入统计,随后才专精于高阶相关性。我们在一个极简数据模型——混合累积量模型(mixed cumulant model)上重现了这一行为,能精确控制输入的成对与高阶相关性。我们识别出一个标量不变量,称为扩散信息指数,类比于其他学习范式中的相关不变量。利用该不变量,我们证明:去噪器以线性样本复杂度学习输入的成对统计,而更复杂的高阶统计(如四阶累积量)至少需要立方样本复杂度。此外,若成对与高阶统计共享相关潜在结构,则四阶累积量的学习复杂度可降至线性。本工作揭示了扩散模型逐步学习递增复杂度分布的关键机制。
原文摘要 · Abstract (English)
While diffusion models have emerged as a powerful class of generative models, their learning dynamics remain poorly understood. We address this issue first by empirically showing that standard diffusion models trained on natural images exhibit a distributional simplicity bias, learning simple, pair-wise input statistics before specializing to higher-order correlations. We reproduce this behaviour in simple denoisers trained on a minimal data model, the mixed cumulant model, where we precisely control both pair-wise and higher-order correlations of the inputs. We identify a scalar invariant of the model that governs the sample complexity of learning pair-wise and higher-order correlations that we call the diffusion information exponent, in analogy to related invariants in different learning paradigms. Using this invariant, we prove that the denoiser learns simple, pair-wise statistics of the inputs at linear sample complexity, while more complex higher-order statistics, such as the fourth cumulant, require at least cubic sample complexity. We also prove that the sample complexity of learning the fourth cumulant is linear if pair-wise and higher-order statistics share a correlated latent structure. Our work describes a key mechanism for how diffusion models can learn distributions of increasing complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。