arXiv:2410.16516cs.LG2024-10被引 6

提出代理方法加速数据记忆评分计算,实现高效可扩展的机器遗忘。

Scalability of memorization-based machine unlearning

  • 用代理指标替代耗时的记忆评分计算,提升算法效率。
  • 在保持高精度和隐私保护的前提下,显著提升可扩展性。
  • 适合需要快速删除敏感数据的实际应用系统使用。

机器遗忘(MUL)旨在消除特定数据子集(如噪声、污染或隐私敏感数据)对预训练模型的影响。现有方法通常依赖专门的微调策略。近期研究发现,数据记忆程度是决定MUL难度的关键特征。因此,基于记忆度的新型遗忘方法被提出,其在遗忘质量与模型效用方面表现优异。然而,这些方法依赖于数据点的记忆评分,而计算这些评分的过程极为耗时,严重限制了其可扩展性与实际应用。本文针对最先进的基于记忆度的MUL算法,采用一系列记忆评分代理指标来解决可扩展性挑战。我们分析了多种代理指标的特性,并评估了先进MUL算法在准确率与隐私保护方面的性能。实验证明,这些代理指标可在保持与完整记忆度方法相当的准确率的同时,大幅提高计算效率。本工作为实现高效、可扩展的机器遗忘迈出了关键一步。

原文摘要 · Abstract (English)

Machine unlearning (MUL) focuses on removing the influence of specific subsets of data (such as noisy, poisoned, or privacy-sensitive data) from pretrained models. MUL methods typically rely on specialized forms of fine-tuning. Recent research has shown that data memorization is a key characteristic defining the difficulty of MUL. As a result, novel memorization-based unlearning methods have been developed, demonstrating exceptional performance with respect to unlearning quality, while maintaining high performance for model utility. Alas, these methods depend on knowing the memorization scores of data points and computing said scores is a notoriously time-consuming process. This in turn severely limits the scalability of these solutions and their practical impact for real-world applications. In this work, we tackle these scalability challenges of state-of-the-art memorization-based MUL algorithms using a series of memorization-score proxies. We first analyze the profiles of various proxies and then evaluate the performance of state-of-the-art (memorization-based) MUL algorithms in terms of both accuracy and privacy preservation. Our empirical results show that these proxies can introduce accuracy on par with full memorization-based unlearning while dramatically improving scalability. We view this work as an important step toward scalable and efficient machine unlearning.

机器遗忘可扩展性数据隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。