arXiv:2607.05898cs.LGcs.CR2026-07

提出可实证检验遗忘算法是否有效的审计工具

Auditing of Unlearning Algorithms

论文配图:Auditing of Unlearning Algorithms
图 1 · 摘自论文原文
  • 用成员推理攻击计算遗忘参数ε的下界
  • 严谨方法ε值极小,经验方法ε值巨大
  • 适用于验证遗忘算法真实效果的实践者

评估遗忘算法是否真正消除训练数据影响仍是一个开放挑战。我们提出一种实用的审计工具,通过成员推理攻击计算数据相关的遗忘参数ε的下界。在多个遗忘算法上的评估发现显著差异:具有严格保证的方法(如模型裁剪和回滚删除)获得极小的ε下界,不违背其遗忘保证;而基于海森矩阵、交替上升下降、在遗忘集上上升及在保留集上微调等经验方法则表现出巨大的ε下界,表明其遗忘效果差。该审计工具基于假设检验框架,可实证驳斥虚假的遗忘声明,并在CIFAR-100和Shakespeare文本数据集上得到验证。

原文摘要 · Abstract (English)

Evaluating whether unlearning algorithms truly remove training data influence remains an open challenge. We propose a practical auditor that computes data-dependent lower bounds on the unlearning parameter $\varepsilon$ using membership inference attacks. Evaluating multiple unlearning algorithms, we find a sharp separation: algorithms with rigorous guarantees, such as model clipping and rewind-to-delete, achieve very small $\varepsilon$ bounds that do not falsify their unlearning guarantees, whereas empirical methods such as Hessian-based unlearning, interleaved ascent-descent, ascent on the forget set, and fine-tuning on the retain set exhibit large bounds, indicating poor unlearning. Our auditor provides a practical tool for empirically falsifying unlearning claims through a hypothesis-testing framework, and we validate it on CIFAR-100 and Shakespeare text.

遗忘学习审计工具成员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。