arXiv:2411.00204cs.CL2024-11中稿 · TMLR 2025被引 2

提出新评估框架,精准测试模型删记忆能力与知识恢复效果。

RESTOR: Knowledge Recovery in Machine Unlearning

  • 设计可同时检测遗忘与知识恢复的评估框架
  • 发现部分算法只强调遗忘而忽略知识重建
  • 适合关注隐私保护与模型可控性的研究者

大规模语言模型在训练过程中可能记忆到包含错误信息、版权内容或敏感个人信息的不良数据。近年来,已有多种机器遗忘算法被提出,旨在消除这些数据对已训练模型的影响——即近似于从未训练过这些数据点的模型状态。然而,评估遗忘算法的有效性仍是开放挑战。以往方法依赖启发式手段,如验证模型是否无法复现特定目标信息,同时保持在无关测试数据上的准确率,但此类方法未能全面捕捉数据点对模型影响的逆转效果。本文提出RESTOR框架,用于机器遗忘评估,通过衡量模型在删除目标数据后遗忘知识的能力,同时评估其能否恢复未曾接触这些数据时的原始知识状态。该框架揭示了若干关于主流遗忘算法的新见解,例如某些算法仅强化遗忘但未实现知识恢复,且局部化遗忘目标能提升性能。

原文摘要 · Abstract (English)

Large language models trained on web-scale corpora can memorize undesirable data containing misinformation, copyrighted material, or private or sensitive information. Recently, several machine unlearning algorithms have been proposed to eliminate the effect of such datapoints from trained models -- that is, to approximate a model that had never been trained on these datapoints in the first place. However, evaluating the effectiveness of unlearning algorithms remains an open challenge. Previous work has relied on heuristics -- such as verifying that the model can no longer reproduce the specific information targeted for removal while maintaining accuracy on unrelated test data. These approaches inadequately capture the complete effect of reversing the influence of datapoints on a trained model. In this work, we propose the RESTOR framework for machine unlearning evaluation, which assesses the ability of unlearning algorithms for targeted data erasure, by evaluating the ability of models to forget the knowledge introduced in these datapoints, while simultaneously recovering the model's knowledge state had it never encountered these datapoints. RESTOR helps uncover several novel insights about popular unlearning algorithms, and the mechanisms through which they operate -- for instance, identifying that some algorithms merely emphasize forgetting but not recovering knowledge, and that localizing unlearning targets can enhance unlearning performance.

机器遗忘知识恢复隐私保护模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。