提出机器遗忘中个体样本难易度差异,揭示遗忘难度的共性影响因素。
Instance-Level Difficulty: A Missing Perspective in Machine Unlearning
- 从样本层面分析遗忘难易度,发现四类通用影响因素。
- 这些因素与具体遗忘算法无关,仅依赖于模型和训练数据。
- 为个性化遗忘策略提供新视角,适合关注隐私保护的研究者。
当前深度机器遗忘研究主要关注整体方法的有效性提升或评估,忽略了单个训练样本遗忘难度的差异。这导致机器遗忘的广泛应用潜力未被充分探索。本文通过在多种遗忘算法和数据集上进行细致的实例级遗忘性能分析,深入研究了使机器遗忘变得困难的核心因素。我们总结出四类使单个数据点难以遗忘的因素,并实证表明这些因素独立于特定遗忘算法,仅与目标模型及其训练数据相关。基于此,我们认为机器遗忘研究应重视个体样本的遗忘难度问题。
原文摘要 · Abstract (English)
Current research on deep machine unlearning primarily focuses on improving or evaluating the overall effectiveness of unlearning methods while overlooking the varying difficulty of unlearning individual training samples. As a result, the broader feasibility of machine unlearning remains under-explored. This paper studies the cruxes that make machine unlearning difficult through a thorough instance-level unlearning performance analysis over various unlearning algorithms and datasets. In particular, we summarize four factors that make unlearning a data point difficult, and we empirically show that these factors are independent of a specific unlearning algorithm but only relevant to the target model and its training data. Given these findings, we argue that machine unlearning research should pay attention to the instance-level difficulty of unlearning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。