arXiv:2410.08557cs.LG2024-10中稿 · Machine Learning J…被引 4

提出可实现精确机器遗忘的理论与算法,突破现有方法局限

MUSO: Achieving Exact Machine Unlearning in Over-Parameterized Regimes

  • 用随机特征构建解析框架,证明过参数化线性模型可通过重标记实现精确遗忘
  • 在真实非线性网络中设计交替优化算法,实验显示性能优于当前最优方法
  • 适合关注数据隐私与模型可控性的研究者,尤其适用于需完全清除特定训练数据的场景

机器遗忘(MU)旨在使训练好的模型表现得像从未学习过特定数据。在当前以神经网络为主导的过参数化模型中,普遍做法是手动重标记数据并微调模型,这可在输出空间近似实现遗忘,但无法保证在参数空间达到精确遗忘。本文通过随机特征技术构建分析框架,在随机梯度下降优化前提下,理论上证明过参数化线性模型可通过重标记特定数据实现精确遗忘。进一步将该方法扩展至真实非线性网络,提出一种统一遗忘与重标记任务的交替优化算法。数值实验验证了该算法在多种场景下的有效性,其性能显著优于现有最先进方法,尤其在同类重标记基方法中表现突出。

原文摘要 · Abstract (English)

Machine unlearning (MU) is to make a well-trained model behave as if it had never been trained on specific data. In today's over-parameterized models, dominated by neural networks, a common approach is to manually relabel data and fine-tune the well-trained model. It can approximate the MU model in the output space, but the question remains whether it can achieve exact MU, i.e., in the parameter space. We answer this question by employing random feature techniques to construct an analytical framework. Under the premise of model optimization via stochastic gradient descent, we theoretically demonstrated that over-parameterized linear models can achieve exact MU through relabeling specific data. We also extend this work to real-world nonlinear networks and propose an alternating optimization algorithm that unifies the tasks of unlearning and relabeling. The algorithm's effectiveness, confirmed through numerical experiments, highlights its superior performance in unlearning across various scenarios compared to current state-of-the-art methods, particularly excelling over similar relabeling-based MU approaches.

机器遗忘过参数化随机特征交替优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。