arXiv:2505.17638cs.LGcond-mat.dis-nn2025-05NeurIPS被引 82

扩散模型不记忆数据,靠训练动态隐式正则化。

Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training

  • 发现训练存在两个时间尺度:生成与记忆的分界点随数据量线性增长。
  • 数据越多,越晚出现记忆,长期训练仍能保持泛化能力。
  • 适合关注生成模型泛化机制的研究者和工程师。

扩散模型在众多生成任务中表现卓越,但其为何不记忆训练数据仍是一个关键问题。本文通过大量实验与理论分析,揭示训练动态中的双重时间尺度:早期τₘₐₚ为模型开始生成高质量样本的时间,后期τₘₑₘ为记忆开始出现的时间。关键发现:τₘₑₘ随训练集大小n线性增加,而τₘₐₚ保持不变。这导致随着n增大,模型拥有更长的泛化窗口,即使继续训练也不会立即过拟合。仅当n超过模型依赖的阈值后,无限训练下过拟合才消失。这一现象源于训练动态中的隐式动力学正则化,使高度过参数化模型也能避免记忆。研究基于标准U-Net在真实与合成数据上的数值实验,以及高维极限下的可解析随机特征模型理论分析,均支持结论。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable success across a wide range of generative tasks. A key challenge is understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time $τ_\mathrm{gen}$ at which models begin to generate high-quality samples, and a later time $τ_\mathrm{mem}$ beyond which memorization emerges. Crucially, we find that $τ_\mathrm{mem}$ increases linearly with the training set size $n$, while $τ_\mathrm{gen}$ remains constant. This creates a growing window of training times with $n$ where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when $n$ becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allow to avoid memorization even in highly overparameterized settings. Our results are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, and by a theoretical analysis using a tractable random features model studied in the high-dimensional limit.

扩散模型泛化能力正则化训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。