arXiv:2503.00592cs.LG2025-03被引 5

提出可评估生成模型记忆单张图像能力的新方法。

SolidMark: Evaluating Image Memorization in Generative Models

  • 设计了针对每张图像的精准记忆评分机制。
  • 能检测到像素级的记忆现象,准确率更高。
  • 适合研究生成模型隐私安全与数据泄露问题的人。

近期研究表明,扩散模型会记忆训练图像并在生成时复现。然而,现有记忆评估指标存在数据集依赖偏差,难以判断特定图像是否被记忆。本文系统分析了扩散模型记忆评估中的问题,提出新方法 $ m ext{SolidMark}$,实现对每张图像的记忆评分。通过重新评估已有缓解技术,验证了该方法的有效性,并展示了其在像素级记忆检测上的能力。最后,开源了基于 $ m ext{SolidMark}$ 的多种模型,以推动对生成模型记忆现象的深入研究。代码已公开于 https://github.com/NickyDCFP/SolidMark。

原文摘要 · Abstract (English)

Recent works have shown that diffusion models are able to memorize training images and emit them at generation time. However, the metrics used to evaluate memorization and its mitigation techniques suffer from dataset-dependent biases and struggle to detect whether a given specific image has been memorized or not. This paper begins with a comprehensive exploration of issues surrounding memorization metrics in diffusion models. Then, to mitigate these issues, we introduce $\rm \style{font-variant: small-caps}{SolidMark}$, a novel evaluation method that provides a per-image memorization score. We then re-evaluate existing memorization mitigation techniques. We also show that $\rm \style{font-variant: small-caps}{SolidMark}$ is capable of evaluating fine-grained pixel-level memorization. Finally, we release a variety of models based on $\rm \style{font-variant: small-caps}{SolidMark}$ to facilitate further research for understanding memorization phenomena in generative models. All of our code is available at https://github.com/NickyDCFP/SolidMark.

生成模型记忆检测扩散模型隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。