arXiv:2504.14798cs.LGcs.CV2025-04被引 2

提出鲁棒遗忘评估框架RUB,检验模型是否真能彻底遗忘敏感信息。

RUB: Evaluating Residual Knowledge in Unlearned Models

  • 设计统一评测框架RUB,检测遗忘后模型残留知识。
  • 现有顶尖遗忘方法在对抗攻击下仍易泄露数据。
  • 适合关注隐私保护与模型安全的研究者使用。

机器遗忘(MUL)作为隐私保护和内容监管的关键机制,当前技术常无法确保敏感信息的完全清除。现有工作多聚焦于验证遗忘执行过程,却忽视了模型是否对多种对抗性攻击具备鲁棒性。本文倡导鲁棒遗忘原则:模型不仅应与重新训练版本难以区分,还须抵御各类对抗威胁。为此,我们提出统一基准RUB(Robust Unlearning Benchmark),系统评估分类、图像到图像重建及文本到图像生成任务中遗忘算法的鲁棒性。在此框架内,引入通用的遗忘映射攻击(UMA)以检测残留信息,并证明现有攻击策略可适配该框架。跨判别与生成任务的实验表明,即使通过标准验证指标,最先进遗忘方法仍易受攻击。通过将鲁棒性作为核心标准并提供对抗评估基准,RUB有望推动更可靠、更安全的遗忘实践。代码库与模型检查点将公开发布。

原文摘要 · Abstract (English)

Machine Unlearning (MUL) has emerged as a key mechanism for privacy protection and content regulation, yet current techniques often fail to guarantee the complete removal of sensitive information. While most existing works focus on verifying the execution of unlearning, they overlook the critical question of whether models remain robust against adversarial attempts to recover forgotten knowledge. In this work, we advocate for the principle of Robust Unlearning, which requires models to be both indistinguishable from retrained counterparts and resilient against diverse adversarial threats. To instantiate this principle, we propose a unified benchmark, RUB (Robust Unlearning Benchmark), that systematically evaluates the robustness of unlearning algorithms across classification, image-to-image reconstruction, and text-to-image synthesis. Within this framework, we introduce the Unlearning Mapping Attack (UMA) as a generalizable method to detect residual information, and demonstrate how existing attack strategies can be adapted into this framework as long as they conform to the generic UMA framework. Our experiments across discriminative and generative tasks reveal that state-of-the-art unlearning methods remain vulnerable under these evaluations, even when passing standard verification metrics. By positioning robustness as the central criterion and providing a benchmark for adversarial evaluation, we hope RUB paves the way toward more reliable and secure unlearning practices. The codebase and model checkpoints in RUB will be published.

机器遗忘隐私保护对抗攻击评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。