提出新评估框架RULI,精准检测模型遗忘漏洞
Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective
- 设计双目标攻击,按样本粒度衡量遗忘效果与隐私风险
- 实测主流遗忘方法在RULI下攻击成功率更高,暴露隐藏风险
- 适用于图像、文本任务,为模型隐私评估提供严格基准
机器遗忘旨在高效移除训练模型中的特定数据,以应对隐私与合规问题。尽管精确遗忘能实现与重新训练等效的数据清除,但在大规模模型中不切实际,因此人们对非精确遗忘方法兴趣日增。然而这些方法缺乏形式化保证,亟需可靠的评估框架来衡量其隐私性与有效性。本文首先指出现有评估框架的若干关键缺陷:侧重平均情况评估、针对随机样本测试、与重新训练基线比较不完整。为此,我们提出RULI(基于似然推理的修正遗忘评估框架),通过双目标攻击,在样本级别同时衡量遗忘效能与隐私风险。实验表明,当前先进遗忘方法在RULI下攻击成功率显著提升,暴露出原有方法低估的隐私隐患。RULI基于博弈论构建,已在图像与文本任务(涵盖分类到生成)上通过实证验证,提供了一种严谨、可扩展且细粒度的遗忘技术评估方法。
原文摘要 · Abstract (English)
Machine unlearning focuses on efficiently removing specific data from trained models, addressing privacy and compliance concerns with reasonable costs. Although exact unlearning ensures complete data removal equivalent to retraining, it is impractical for large-scale models, leading to growing interest in inexact unlearning methods. However, the lack of formal guarantees in these methods necessitates the need for robust evaluation frameworks to assess their privacy and effectiveness. In this work, we first identify several key pitfalls of the existing unlearning evaluation frameworks, e.g., focusing on average-case evaluation or targeting random samples for evaluation, incomplete comparisons with the retraining baseline. Then, we propose RULI (Rectified Unlearning Evaluation Framework via Likelihood Inference), a novel framework to address critical gaps in the evaluation of inexact unlearning methods. RULI introduces a dual-objective attack to measure both unlearning efficacy and privacy risks at a per-sample granularity. Our findings reveal significant vulnerabilities in state-of-the-art unlearning methods, where RULI achieves higher attack success rates, exposing privacy risks underestimated by existing methods. Built on a game-based foundation and validated through empirical evaluations on both image and text data (spanning tasks from classification to generation), RULI provides a rigorous, scalable, and fine-grained methodology for evaluating unlearning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。