提出按遗忘难易排序的机器遗忘框架,提升隐私数据清除效率。
Machine Unlearning in Forgettability Sequence
- 根据隐私风险高低对数据排序,分步遗忘更精准
- 新框架在保持模型性能前提下,显著提升遗忘成功率
- 适合需要高精度数据清除的合规场景
机器遗忘(MU)正成为实现“被遗忘权”的有前景范式,可消除特定数据点的训练痕迹,同时保持模型在通用测试样本上的性能。随着遗忘研究进展,诸多基础问题仍待解答:不同样本的遗忘难度是否不同?遗忘顺序是否影响算法表现?本文识别出影响遗忘难易的关键因素:高隐私风险样本更易被遗忘,表明遗忘难度存在差异,从而催生更精细的遗忘模式。基于此,提出通用遗忘框架RSU,包含排序模块(Ranking)与序列遗忘模块(SeqUnlearn)。
原文摘要 · Abstract (English)
Machine unlearning (MU) is becoming a promising paradigm to achieve the "right to be forgotten", where the training trace of any chosen data points could be eliminated, while maintaining the model utility on general testing samples after unlearning. With the advancement of forgetting research, many fundamental open questions remain unanswered: do different samples exhibit varying levels of difficulty in being forgotten? Further, does the sequence in which samples are forgotten, determined by their respective difficulty levels, influence the performance of forgetting algorithms? In this paper, we identify key factor affecting unlearning difficulty and the performance of unlearning algorithms. We find that samples with higher privacy risks are more likely to be unlearning, indicating that the unlearning difficulty varies among different samples which motives a more precise unlearning mode. Built upon this insight, we propose a general unlearning framework, dubbed RSU, which consists of Ranking module and SeqUnlearn module.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。