揭示扩散模型记忆与泛化之间的理论差异,提出抑制记忆的新方法。
Provable Separations between Memorization and Generalization in Diffusion Models
- 从统计估计和网络逼近双视角证明记忆与泛化的本质分离
- 实证得分函数需网络规模随样本量增长,而真实得分函数更紧凑
- 基于剪枝的改进方法在降低记忆的同时保持生成质量
扩散模型在多个领域取得显著成功,但依然存在记忆问题——即复制训练数据而非生成新内容。这不仅限制其创造力,也引发隐私与安全担忧。尽管已有大量实证研究探索缓解策略,但对记忆现象的理论理解仍不充分。本文通过统计估计与网络逼近两个互补视角,建立双重分离结果:一方面,真实得分函数并不最小化经验去噪损失,导致记忆倾向;另一方面,实现经验得分函数所需的网络规模必须随样本量增长,远大于真实得分函数的紧凑表示。基于这些洞察,我们提出一种基于剪枝的方法,在保持生成质量的同时有效减少记忆现象。
原文摘要 · Abstract (English)
Diffusion models have achieved remarkable success across diverse domains, but they remain vulnerable to memorization -- reproducing training data rather than generating novel outputs. This not only limits their creative potential but also raises concerns about privacy and safety. While empirical studies have explored mitigation strategies, theoretical understanding of memorization remains limited. We address this gap through developing a dual-separation result via two complementary perspectives: statistical estimation and network approximation. From the estimation side, we show that the ground-truth score function does not minimize the empirical denoising loss, creating a separation that drives memorization. From the approximation side, we prove that implementing the empirical score function requires network size to scale with sample size, spelling a separation compared to the more compact network representation of the ground-truth score function. Guided by these insights, we develop a pruning-based method that reduces memorization while maintaining generation quality in diffusion transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。