提出新指标评估模型遗忘效果,提升数据清除可靠性。
Efficient Unlearning through Maximizing Relearning Convergence Delay

- 用重训练收敛延迟衡量模型对遗忘数据的理解程度
- 方法使遗忘数据恢复难度显著增加,保留数据性能稳定
- 适合需要安全删除敏感数据的场景
机器遗忘在移除错误标注、污染或有问题的数据方面面临挑战。现有方法和评估指标仅关注模型预测,难以反映模型对数据本质特征的真实理解。为此,我们引入新的度量指标——重训练收敛延迟,该指标同时捕捉权重空间与预测空间的变化,提供更全面的模型对遗忘数据的理解评估,可用于判断遗忘数据是否可能从已遗忘模型中被恢复。基于此,我们提出影响消除遗忘框架(Influence Eliminating Unlearning),通过降低遗忘集性能、引入权重衰减和权重噪声来消除其影响,同时保持对保留集的准确率。大量实验表明,该方法优于现有指标及所提收敛延迟度量,逼近理想遗忘性能。我们提供了理论保证,包括指数收敛性与上界分析,以及在分类与生成式遗忘任务中的强保留性和抗重训练证据。
原文摘要 · Abstract (English)
Machine unlearning poses challenges in removing mislabeled, contaminated, or problematic data from a pretrained model. Current unlearning approaches and evaluation metrics are solely focused on model predictions, which limits insight into the model's true underlying data characteristics. To address this issue, we introduce a new metric called relearning convergence delay, which captures both changes in weight space and prediction space, providing a more comprehensive assessment of the model's understanding of the forgotten dataset. This metric can be used to assess the risk of forgotten data being recovered from the unlearned model. Based on this, we propose the Influence Eliminating Unlearning framework, which removes the influence of the forgetting set by degrading its performance and incorporates weight decay and injecting noise into the model's weights, while maintaining accuracy on the retaining set. Extensive experiments show that our method outperforms existing metrics and our proposed relearning convergence delay metric, approaching ideal unlearning performance. We provide theoretical guarantees, including exponential convergence and upper bounds, as well as empirical evidence of strong retention and resistance to relearning in both classification and generative unlearning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。