arXiv:2605.22050cs.CV2026-05KDD

发现扩散模型生成破损图像源于记忆过拟合,提出实时检测与修复方法。

Broken Memories: Detecting and Mitigating Memorization in Diffusion Models with Degraded Generations

论文配图:Broken Memories: Detecting and Mitigating Memorization in Diffusion Models with Degraded Generations
图 1 · 摘自论文原文
  • 基于潜在空间更新范数构建稳定性区域,量化生成过程稳定性。
  • 在Stable Diffusion 1.4上实现>0.999的检测AUC和0%记忆率。
  • 无需改提示词或引导,零成本修复,适合高安全需求场景。

尽管扩散模型能生成高质量图像,但其对训练数据的记忆倾向带来了严重的隐私与版权风险。本文首次发现,记忆行为会引发内部数值不稳定性,常表现为视觉上的“破碎”伪影。受数值方法稳定性分析启发,我们基于潜在更新范数提出经验稳定性区域,定量刻画生成过程中的稳定行为。据此,我们设计了一种原理性、可实时执行的分步检测与自适应缓解框架。该方法在不修改提示词或引导策略的前提下抑制记忆现象,同时保持语义一致性和图像质量。在Stable Diffusion 1.4上的大量实验表明,该方法检测性能达AUC >0.999,缓解后记忆率为0.0%,且计算开销极低(每图约0.01秒)。

原文摘要 · Abstract (English)

While diffusion models excel at generating high-quality images, their tendency to memorize training data poses significant privacy and copyright risks. In this work, we for the first time identify that memorization induces internal numerical instability, often manifesting as visually ``broken'' artifacts. Inspired by stability analysis in numerical methods, we introduce empirical stability regions based on latent update norms to quantitatively characterize stable behavior during generation. Leveraging this, we propose a principled, on-the-fly framework for step-wise detection and adaptive mitigation. Our approach suppresses memorization without altering prompts or guidance, thereby preserving semantic fidelity and image quality. Extensive experiments on Stable Diffusion 1.4 demonstrate that our method achieves an AUC $>0.999$ detection performance and a $0.0\%$ memorization rate after mitigation with negligible overhead ($\approx0.01$s per image).

扩散模型记忆检测图像安全稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。