研究相似样本对模型遗忘的影响,发现现有方法效果不佳
Forgetting Similar Samples: Can Machine Unlearning Do it Better?
- 对比分析多种遗忘方法在存在相似样本时的表现
- 实验显示多数方法无法彻底消除目标样本影响
- 适用于关注模型安全与数据可控性的研究人员
机器遗忘旨在让预训练模型消除特定训练样本的影响,近年来备受关注。尽管已有大量研究致力于提升遗忘效率,但我们认为这些方法主要关注移除样本本身,而非彻底消除其对模型的影响,偏离了机器遗忘的本质定义。本文首次系统评估现有遗忘方案在训练集包含大量与目标样本相似数据时的有效性。通过在四个精心构建的数据集上进行详尽实验与分析,结果表明:即使采用从头再训练的基线方法,多数现有遗忘技术在图像和语言模型上的实际表现仍显著低于预期。此外,本文还探讨了改进当前遗忘方法的潜在方向。
原文摘要 · Abstract (English)
Machine unlearning, a process enabling pre-trained models to remove the influence of specific training samples, has attracted significant attention in recent years. Although extensive research has focused on developing efficient machine unlearning strategies, we argue that these methods mainly aim at removing samples rather than removing samples' influence on the model, thus overlooking the fundamental definition of machine unlearning. In this paper, we first conduct a comprehensive study to evaluate the effectiveness of existing unlearning schemes when the training dataset includes many samples similar to those targeted for unlearning. Specifically, we evaluate: Do existing unlearning methods truly adhere to the original definition of machine unlearning and effectively eliminate all influence of target samples when similar samples are present in the training dataset? Our extensive experiments, conducted on four carefully constructed datasets with thorough analysis, reveal a notable gap between the expected and actual performance of most existing unlearning methods for image and language models, even for the retraining-from-scratch baseline. Additionally, we also explore potential solutions to enhance current unlearning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。