arXiv:2604.05669stat.MLcs.LG2026-04

提出高效机器遗忘方法,实现近似重训练效果且无需全部数据。

Efficient machine unlearning with minimax optimality

  • 基于通用损失函数构建统计框架,针对平方损失设计最优遗忘算法。
  • 在仅用少量剩余数据和遗忘样本时,误差逼近理论下限。
  • 无需全量重训练即可完成可信推断,适合隐私合规与数据净化场景。

随着GDPR等法规对数据删除需求的增长,机器遗忘成为重要研究方向,旨在消除特定数据影响而无需全量重训练。本文提出一种适用于通用损失函数的统计框架,并针对平方损失设计了遗忘最小二乘法(ULS),在仅有预训练估计器、遗忘样本及少量剩余数据的情况下,建立了模型参数估计的极小极大最优性。结果表明,估计误差可分解为一个理想项与由遗忘比例和遗忘模型偏差决定的遗忘代价项。进一步建立了无需全量重训练的渐近有效推断方法。数值实验与真实数据应用显示,该方法性能接近重训练,但数据访问量显著减少。

原文摘要 · Abstract (English)

There is a growing demand for efficient data removal to comply with regulations like the GDPR and to mitigate the influence of biased or corrupted data. This has motivated the field of machine unlearning, which aims to eliminate the influence of specific data subsets without the cost of full retraining. In this work, we propose a statistical framework for machine unlearning with generic loss functions and establish theoretical guarantees. For squared loss, especially, we develop Unlearning Least Squares (ULS) and establish its minimax optimality for estimating the model parameter of remaining data when only the pre-trained estimator, forget samples, and a small subsample of the remaining data are available. Our results reveal that the estimation error decomposes into an oracle term and an unlearning cost determined by the forget proportion and the forget model bias. We further establish asymptotically valid inference procedures without requiring full retraining. Numerical experiments and real-data applications demonstrate that the proposed method achieves performance close to retraining while requiring substantially less data access.

机器遗忘统计学习数据合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。