提出可逆遗忘攻击与新型记忆替代遗忘方法,提升扩散模型隐私安全性
Towards Irreversible Machine Unlearning for Diffusion Models
- 设计攻击方法DiMRA,能逆转基于微调的遗忘技术
- 实验验证现有遗忘方法易被逆向,生成被删除内容
- 提出DiMUM方法,通过记忆替代实现更安全的遗忘
扩散模型在生成合成图像方面表现卓越,但其带来的安全、隐私和版权问题促使机器遗忘技术的发展,以使模型忘记特定训练数据并防止生成敏感内容。当前针对扩散模型的遗忘方法主要面向条件扩散模型,聚焦于特定类别或特征的遗忘,其中基于微调的方法因效率高、效果好而被广泛采用,通过最小化精心设计的损失函数来更新预训练模型参数。然而本文提出一种新型攻击——扩散模型重学攻击(DiMRA),可逆向这类基于微调的遗忘方法,且无需事先知晓被遗忘元素即可通过辅助数据优化被遗忘模型,恢复出已删除的内容。为应对该漏洞,本文提出一种新遗忘方法——扩散模型记忆式遗忘(DiMUM),不同于传统遗忘思路,DiMUM不直接抹除目标数据或特征,而是通过记忆替代数据或特征来阻止生成原目标内容。实验表明,DiMRA能有效逆转现有先进微调式遗忘方法,凸显了对更鲁棒解决方案的需求;全面评估显示,DiMUM在保持生成性能的同时显著提升了对抗DiMRA的鲁棒性。
原文摘要 · Abstract (English)
Diffusion models are renowned for their state-of-the-art performance in generating synthetic images. However, concerns related to safety, privacy, and copyright highlight the need for machine unlearning, which can make diffusion models forget specific training data and prevent the generation of sensitive or unwanted content. Current machine unlearning methods for diffusion models are primarily designed for conditional diffusion models and focus on unlearning specific data classes or features. Among these methods, finetuning-based machine unlearning methods are recognized for their efficiency and effectiveness, which update the parameters of pre-trained diffusion models by minimizing carefully designed loss functions. However, in this paper, we propose a novel attack named Diffusion Model Relearning Attack (DiMRA), which can reverse the finetuning-based machine unlearning methods, posing a significant vulnerability of this kind of technique. Without prior knowledge of the unlearning elements, DiMRA optimizes the unlearned diffusion model on an auxiliary dataset to reverse the unlearning, enabling the model to regenerate previously unlearned elements. To mitigate this vulnerability, we propose a novel machine unlearning method for diffusion models, termed as Diffusion Model Unlearning by Memorization (DiMUM). Unlike traditional methods that focus on forgetting, DiMUM memorizes alternative data or features to replace targeted unlearning data or features in order to prevent generating such elements. In our experiments, we demonstrate the effectiveness of DiMRA in reversing state-of-the-art finetuning-based machine unlearning methods for diffusion models, highlighting the need for more robust solutions. We extensively evaluate DiMUM, demonstrating its superior ability to preserve the generative performance of diffusion models while enhancing robustness against DiMRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。