arXiv:2606.29832cs.LG2026-06

破解持续学习中遗忘与保留的难题,提出首个理论框架。

The Forgetting-Retention Dilemma: Certified Unlearning Theory in Continual Learning

论文配图:The Forgetting-Retention Dilemma: Certified Unlearning Theory in Continual Learning
图 1 · 摘自论文原文
  • 构建持续学习下可认证遗忘的理论模型,明确遗忘与知识保留的权衡。
  • 证明梯度法虽遗忘效果弱,但存储开销近乎为零,优于海森法。
  • 提出混合策略,在低存储成本下保持良好遗忘性能,适合隐私敏感场景。

机器遗忘旨在消除特定数据对训练模型的影响以保护隐私。然而,在持续学习(CL)框架中,模型需在动态数据集上顺序更新,当前可认证遗忘算法未能考虑这种复杂的累积模型演化过程。本文首次建立连接持续学习与机器遗忘的理论基础。我们将持续学习的遗忘目标定义为最小化遗忘后的额外风险,该风险由持续学习额外风险和遗忘损失构成,刻画了保留历史知识与目标遗忘之间的根本权衡。在温和假设下,我们首先建立了非凸模型中持续学习额外风险的上界。随后,将两种可认证遗忘方法——基于梯度和基于海森的方法,适配至持续学习框架。分析表明,尽管基于梯度的方法在最小化遗忘损失方面不如基于海森的方法,但其具有近乎零的存储开销优势,便于实现遗忘。这一发现启发我们提出一种混合策略,在降低存储成本的同时维持遗忘后性能。实验结果进一步验证了理论结论。

原文摘要 · Abstract (English)

Machine unlearning aims to eliminate the influence of specific data from trained models to safeguard privacy. However, this presents a significant challenge in the context of continual learning (CL), where models update sequentially on dynamic datasets. A major limitation is that current certified unlearning algorithms fail to account for the complex, cumulative model evolution inherent to CL framework. In this work, we establish the first theoretical foundation bridging CL and machine unlearning. We formulate the CL's unlearning objective as the minimization of post-unlearning excess risk, which decomposes into CL excess risk and unlearning loss, characterizing the fundamental trade-off between preserving historical knowledge and targeted forgetting. Under mild assumptions, we first establish an upper bound for the CL excess risk in non-convex models. We then adapt two certified unlearning approaches, gradient-based and Hessian-based, to the CL framework. Our analysis reveals that while the gradient-based approach is less effective than the Hessian-based method in minimizing unlearning loss, it offers the distinct advantage of nearly zero storage overhead for enabling unlearning. This insight motivates a hybrid strategy that reduces storage costs while maintaining post-unlearning performance. Experimental results further validate our theoretical findings.

持续学习机器遗忘理论分析隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。