arXiv:2505.07026cs.LGstat.ML2025-05被引 1

提出高效且可验证的机器遗忘方法,支持精确删除数据而无需重新训练。

Efficient Machine Unlearning by Model Splitting and Core Sample Selection

  • 通过模型拆分与核心样本选择实现高效遗忘
  • 在多数情况下可实现精确遗忘,性能接近全量重训练
  • 适用于需满足数据删除法律要求的场景

机器遗忘对于满足法律义务(如被遗忘权)至关重要,要求在请求时从机器学习模型中移除特定数据。尽管已有多种遗忘方法,但现有方案通常效率低下,且难以验证遗忘效果——尤其在弱遗忘保证下,验证仍是开放挑战。本文提出一种标准遗忘指标的广义变体,使遗忘策略更高效、更精确。同时引入一种感知遗忘的训练过程,在许多情况下可实现精确遗忘。该方法称为MaxRR。当精确遗忘不可行时,MaxRR仍能以接近全量重训练的性能实现高效遗忘。

原文摘要 · Abstract (English)

Machine unlearning is essential for meeting legal obligations such as the right to be forgotten, which requires the removal of specific data from machine learning models upon request. While several approaches to unlearning have been proposed, existing solutions often struggle with efficiency and, more critically, with the verification of unlearning - particularly in the case of weak unlearning guarantees, where verification remains an open challenge. We introduce a generalized variant of the standard unlearning metric that enables more efficient and precise unlearning strategies. We also present an unlearning-aware training procedure that, in many cases, allows for exact unlearning. We term our approach MaxRR. When exact unlearning is not feasible, MaxRR still supports efficient unlearning with properties closely matching those achieved through full retraining.

机器遗忘数据删除模型更新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。