提出新方法保护图像生成模型在删训后仍保持完整性。
Model Integrity when Unlearning with T2I Diffusion Models
- 设计新评估指标,直接衡量删训前后生成图像的感知差异。
- 新算法在保留生成能力上显著优于现有方法。
- 适合关注模型安全与隐私保护的研究者使用。
文本到图像扩散模型的快速发展使其广泛可得,但这些模型因训练数据来自互联网,可能生成不当内容。为缓解此问题,已有近似机器删训算法通过调整模型权重,降低特定类型图像(来自'遗忘分布')的生成概率,同时尽量保留其他图像(来自'保留分布')的生成能力。然而,我们指出这类方法可能损害模型完整性,意外影响保留分布图像的生成。针对FID和CLIPScore在捕捉此类影响上的局限性,本文提出一种新的保留度量,直接评估原模型与删训后模型生成结果间的感知差异。基于该度量,我们设计的新删训算法在保持模型完整性方面显著优于现有基线。由于实现简单,这些算法可作为未来扩散模型近似机器删训研究的重要基准。
原文摘要 · Abstract (English)
The rapid advancement of text-to-image Diffusion Models has led to their widespread public accessibility. However these models, trained on large internet datasets, can sometimes generate undesirable outputs. To mitigate this, approximate Machine Unlearning algorithms have been proposed to modify model weights to reduce the generation of specific types of images, characterized by samples from a ``forget distribution'', while preserving the model's ability to generate other images, characterized by samples from a ``retain distribution''. While these methods aim to minimize the influence of training data in the forget distribution without extensive additional computation, we point out that they can compromise the model's integrity by inadvertently affecting generation for images in the retain distribution. Recognizing the limitations of FID and CLIPScore in capturing these effects, we introduce a novel retention metric that directly assesses the perceptual difference between outputs generated by the original and the unlearned models. We then propose unlearning algorithms that demonstrate superior effectiveness in preserving model integrity compared to existing baselines. Given their straightforward implementation, these algorithms serve as valuable benchmarks for future advancements in approximate Machine Unlearning for Diffusion Models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。