arXiv:2501.15785cs.LGmath.DS2025-01被引 41

扩散模型易记忆训练数据,本文揭示其机制并提出正则化方案。

Memorization and Regularization in Generative Diffusion Models

  • 发现扩散模型在经验损失下会收敛到高斯混合得分函数
  • 实验验证多种正则化可抑制对训练样本的直接复制
  • 适合关注生成模型泛化能力的研究者阅读

扩散模型作为生成建模的强大框架,其核心是分数匹配:学习不同尺度噪声数据分布的对数密度梯度。当分数匹配使用经验数据而非总体损失时,最优解对应于一个时变高斯混合的得分函数。然而,采用这一解析可解的最小值会导致数据记忆:在无条件和有条件设置下,生成模型均返回训练样本。本文分析了记忆现象的动力学机制,强调正则化对避免重现该解析解的必要性,并为正则化设计提供理论基础。数值实验研究了三种正则化方法:(i) Tikhonov 正则化;(ii) 促进渐近一致性的正则化;(iii) 由神经网络欠参数化或早停引起的正则化。这些实验在记忆现象背景下评估,指出了未来正则化发展的方向。

原文摘要 · Abstract (English)

Diffusion models have emerged as a powerful framework for generative modeling. At the heart of the methodology is score matching: learning gradients of families of log-densities for noisy versions of the data distribution at different scales. When the loss function adopted in score matching is evaluated using empirical data, rather than the population loss, the minimizer corresponds to the score of a time-dependent Gaussian mixture. However, use of this analytically tractable minimizer leads to data memorization: in both unconditioned and conditioned settings, the generative model returns the training samples. This paper contains an analysis of the dynamical mechanism underlying memorization. The analysis highlights the need for regularization to avoid reproducing the analytically tractable minimizer; and, in so doing, lays the foundations for a principled understanding of how to regularize. Numerical experiments investigate the properties of: (i) Tikhonov regularization; (ii) regularization designed to promote asymptotic consistency; and (iii) regularizations induced by under-parameterization of a neural network or by early stopping when training a neural network. These experiments are evaluated in the context of memorization, and directions for future development of regularization are highlighted.

扩散模型正则化数据记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。