扩散模型为何不记忆数据却能生成新样本?需重新思考其泛化机制。
Understanding diffusion models requires rethinking (again) generalization

- 从记忆与泛化的矛盾出发,提出应关注预记忆阶段模型学到了什么
- 实验证明模型在完全记忆前已开始学习数据分布本质特征
- 适合研究生成模型理论机制的学者参考
本文主张,理解扩散模型的泛化能力需要全新的理论框架,不能仅依赖传统统计学习理论或监督学习中的良性过拟合范式。在扩散模型中,训练数据的完全记忆与生成新样本的能力相互排斥:一旦完全记忆,模型只会复制已有数据而非生成新内容。尽管已有研究尝试用容量限制、优化过程中的隐式正则化或架构归纳偏置解释实际扩散模型为何仍能泛化,但这些因素的交互机制尚不明确。我们提出,研究重点应从解释为何扩散模型不记忆,转向探究其在记忆发生前究竟学到了什么。为此,我们在CIFAR-10上进行实证研究,并提炼出若干关键开放问题,以推动对扩散模型泛化机制的深入理解。
原文摘要 · Abstract (English)
This position paper argues that understanding generalization in diffusion models requires fundamentally new theoretical frameworks that go beyond both classical statistical learning theory and the benign overfitting paradigm developed for supervised learning. In diffusion models, unlike in supervised learning, memorization of training data and generalization to novel samples are incompatible: a model that has fully memorized its training set generates copies rather than novel data. Several theoretical explanations for why practical diffusion models nevertheless generalize have been proposed, based on capacity limitations, implicit regularization from optimization, or architectural inductive biases, but their interactions remain unclear. We argue that the field should pivot from explaining why the diffusion models do not memorize to investigating what the model actually learns during pre-memorization phase. To highlight our stance, we conduct empirical study of diffusion models trained on CIFAR-10, and we distill the findings into concrete open questions that we believe are key to improve understanding of generalization in diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。