揭示扩散模型记忆训练数据的核心机制
How Diffusion Models Memorize
- 早期去噪阶段对训练样本的过估计是记忆主因
- 记忆导致潜空间轨迹坍缩,加速收敛到特定图像
- 适合关注隐私与模型安全的研究者阅读
尽管扩散模型在图像生成中表现优异,但其可能记忆训练数据,引发严重隐私和版权问题。现有研究试图描述、检测和缓解记忆现象,但其根本原因仍不明确。本文重新审视扩散与去噪过程,分析潜空间动态,回答‘扩散模型如何记忆?’这一问题。研究发现,记忆源于早期去噪阶段对训练样本的过估计,这降低了多样性,使去噪轨迹坍缩,并加速收敛至记忆图像。具体而言:(i) 记忆不能仅用过拟合解释,因在记忆状态下训练损失更大,这是由于无分类器引导放大预测并引发过估计;(ii) 记忆提示会将训练图像注入噪声预测,迫使潜空间轨迹收敛并引导去噪朝向对应样本;(iii) 中间潜变量分解显示初始随机性迅速被记忆内容取代,且偏离理论去噪调度的程度与记忆严重性几乎完全相关。这些结果共同指出,早期过估计是扩散模型记忆的核心机制。
原文摘要 · Abstract (English)
Despite their success in image generation, diffusion models can memorize training data, raising serious privacy and copyright concerns. Although prior work has sought to characterize, detect, and mitigate memorization, the fundamental question of why and how it occurs remains unresolved. In this paper, we revisit the diffusion and denoising process and analyze latent space dynamics to address the question: "How do diffusion models memorize?" We show that memorization is driven by the overestimation of training samples during early denoising, which reduces diversity, collapses denoising trajectories, and accelerates convergence toward the memorized image. Specifically: (i) memorization cannot be explained by overfitting alone, as training loss is larger under memorization due to classifier-free guidance amplifying predictions and inducing overestimation; (ii) memorized prompts inject training images into noise predictions, forcing latent trajectories to converge and steering denoising toward their paired samples; and (iii) a decomposition of intermediate latents reveals how initial randomness is quickly suppressed and replaced by memorized content, with deviations from the theoretical denoising schedule correlating almost perfectly with memorization severity. Together, these results identify early overestimation as the central underlying mechanism of memorization in diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。