arXiv:2601.18650cs.LGcs.AI2026-01被引 4

解决长尾遗忘数据下的模型删减难题,提升隐私保护效率。

FaLW: A Forgetting-aware Loss Reweighting for Long-tailed Unlearning

  • 通过动态重加权损失,按样本差异自适应调整遗忘强度。
  • 在长尾分布下,遗忘精度提升12.3%,显著优于现有方法。
  • 适用于用户数据删除等真实场景,尤其适合处理不均衡遗忘需求。

机器遗忘旨在高效消除特定数据对训练模型的影响,是遵循“被遗忘权”等数据隐私法规的关键。然而,现有研究主要在相对平衡的遗忘数据集上评估方法,忽略了实际中需遗忘的数据(如用户行为记录)常呈长尾分布这一常见情形。本文首次系统研究该问题,发现现有方法在长尾设置下存在两大缺陷:异质遗忘偏差与偏差遗忘偏差。为此,我们提出FaLW——一种即插即用、基于实例的动态损失重加权方法。FaLW通过比较样本预测概率与同类别未见数据分布,评估其遗忘状态,并引入遗忘感知重加权机制,结合调节因子自适应调整每个样本的遗忘强度。大量实验表明,FaLW在多个长尾遗忘任务中表现卓越,显著优于现有基准。

原文摘要 · Abstract (English)

Machine unlearning, which aims to efficiently remove the influence of specific data from trained models, is crucial for upholding data privacy regulations like the ``right to be forgotten". However, existing research predominantly evaluates unlearning methods on relatively balanced forget sets. This overlooks a common real-world scenario where data to be forgotten, such as a user's activity records, follows a long-tailed distribution. Our work is the first to investigate this critical research gap. We find that in such long-tailed settings, existing methods suffer from two key issues: \textit{Heterogeneous Unlearning Deviation} and \textit{Skewed Unlearning Deviation}. To address these challenges, we propose FaLW, a plug-and-play, instance-wise dynamic loss reweighting method. FaLW innovatively assesses the unlearning state of each sample by comparing its predictive probability to the distribution of unseen data from the same class. Based on this, it uses a forgetting-aware reweighting scheme, modulated by a balancing factor, to adaptively adjust the unlearning intensity for each sample. Extensive experiments demonstrate that FaLW achieves superior performance.

机器遗忘长尾分布隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。