arXiv:2607.02671stat.MLcs.LG2026-07

扩散模型中不存在良性过拟合,过拟合会损害泛化性能。

Benign Overfitting Does Not Occur in Diffusion Models

论文配图:Benign Overfitting Does Not Occur in Diffusion Models
图 1 · 摘自论文原文
  • 证明在数据维度高时,过拟合与良好泛化无法共存。
  • 损失曲线呈经典U形而非双下降,模型复杂度过高会恶化性能。
  • 时间平滑性和早停机制是防止过拟合的关键,适合研究生成模型理论者阅读。

良性过拟合和双下降现象重塑了我们对深度学习泛化的理解,表明过拟合不仅可与良好泛化共存,甚至能促进其效果。扩散模型虽共享传统深度学习的大部分架构,但本工作揭示这一假设大多不成立。我们首先建立根本性不可能结果:除非样本量随数据维度指数增长,否则过拟合与良好泛化无法同时实现。因此,总体损失随模型复杂度呈现经典U形曲线,而非双下降。在简化设置下,我们发现回归依赖目标与经验协方差对齐以获益,而得分匹配无此对齐机制,导致过拟合不可逆地有害。此外,我们识别出得分的时间平滑性及训练中的早停机制构成隐式正则化,有效抑制过拟合,并通过高维图像生成实验验证结论。结果表明,扩散模型的泛化机制与传统回归显著不同,亟需发展新理论。

原文摘要 · Abstract (English)

Benign overfitting and double descent have come to shape our understanding of generalization in deep learning, establishing that overfitting is not only compatible with good generalization but can actively benefit it. Diffusion models share much of the machinery of standard deep learning, so it is natural to assume that they also exhibit these properties. In this work, we show that this assumption is largely incorrect. We first establish fundamental impossibility results showing that, unless the sample size grows exponentially with the data dimension, overfitting and good generalization cannot occur simultaneously. Consequently, the population loss follows a classical U-shaped curve in model complexity rather than exhibiting double descent. Analyzing a simplified setting, we identify a key difference between regression and score matching: regression benefits from an alignment between the target and the empirical covariance; score matching admits no such alignment, leaving overfitting irreparably harmful. We further identify implicit regularization stemming from time-smoothness of the score and early stopping during training as mechanisms that prevent such overfitting and verify our findings with high-dimensional image generation experiments. Our results reveal that generalization in diffusion models is governed by mechanisms distinct from those of traditional regression, motivating the development of new theory.

扩散模型泛化理论过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。