提出记忆自再生机制,让被遗忘模型重新找回知识
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
- 设计MemoRa策略实现被遗忘知识的恢复
- 发现遗忘分短期可快速召回与长期难恢复两种模式
- 强调知识检索鲁棒性是评估遗忘技术的关键
现代文生图模型虽能生成逼真图像,却易被用于制造有害、虚假或非法内容,推动机器遗忘技术发展。该领域旨在选择性移除特定知识,同时保持模型整体性能。然而,真正遗忘某概念极为困难:模型在对抗提示攻击下仍能生成被遗忘内容,可能违法且有害。本文探讨模型遗忘与回忆能力,提出记忆自再生任务;提出MemoRa策略,支持已丢失知识的有效恢复;并主张知识检索鲁棒性是评估遗忘技术的重要指标。实验表明遗忘分两种:短期可快速召回,长期则更难恢复。代码已开源。
原文摘要 · Abstract (English)
The impressive capability of modern text-to-image models to generate realistic visuals has come with a serious drawback: they can be misused to create harmful, deceptive or unlawful content. This has accelerated the push for machine unlearning. This new field seeks to selectively remove specific knowledge from a model's training data without causing a drop in its overall performance. However, it turns out that actually forgetting a given concept is an extremely difficult task. Models exposed to attacks using adversarial prompts show the ability to generate so-called unlearned concepts, which can be not only harmful but also illegal. In this paper, we present considerations regarding the ability of models to forget and recall knowledge, introducing the Memory Self-Regeneration task. Furthermore, we present MemoRa strategy, which we consider to be a regenerative approach supporting the effective recovery of previously lost knowledge. Moreover, we propose that robustness in knowledge retrieval is a crucial yet underexplored evaluation measure for developing more robust and effective unlearning techniques. Finally, we demonstrate that forgetting occurs in two distinct ways: short-term, where concepts can be quickly recalled, and long-term, where recovery is more challenging. Code is available at https://gmum.github.io/MemoRa/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。