arXiv:2505.23270cs.LGcs.AI2025-05被引 6

测试大模型删知识是否真能删干净,发现现有方法有漏洞

Does Machine Unlearning Truly Remove Knowledge?

  • 设计三套数据集+六种删除算法+五种检测方法,系统评估删知识效果
  • 实测显示多数方法无法彻底清除敏感信息,存在知识残留
  • 提出新检测法:扰动中间激活值,比只看输入输出更敏感

近年来,大语言模型在大规模架构和海量数据训练下取得显著进展,但其训练数据常包含来自公共网络的敏感或受版权保护内容,引发数据隐私与所有权担忧。根据通用数据保护条例(GDPR)等法规,个人有权要求移除其个人信息。这推动了机器遗忘算法的发展,旨在不重新训练的前提下移除模型中的特定知识。然而,由于大模型的复杂性和生成特性,评估遗忘算法的有效性仍具挑战。本文提出一个全面的遗忘评估审计框架,包含三个基准数据集、六种遗忘算法和五种基于提示的审计方法。通过多种审计算法,评估不同遗忘策略的有效性与鲁棒性。为进一步探索输入输出依赖型审计的局限,提出一种基于中间激活值扰动的新技术,以提升检测敏感信息残留的能力。

原文摘要 · Abstract (English)

In recent years, Large Language Models (LLMs) have achieved remarkable advancements, drawing significant attention from the research community. Their capabilities are largely attributed to large-scale architectures, which require extensive training on massive datasets. However, such datasets often contain sensitive or copyrighted content sourced from the public internet, raising concerns about data privacy and ownership. Regulatory frameworks, such as the General Data Protection Regulation (GDPR), grant individuals the right to request the removal of such sensitive information. This has motivated the development of machine unlearning algorithms that aim to remove specific knowledge from models without the need for costly retraining. Despite these advancements, evaluating the efficacy of unlearning algorithms remains a challenge due to the inherent complexity and generative nature of LLMs. In this work, we introduce a comprehensive auditing framework for unlearning evaluation, comprising three benchmark datasets, six unlearning algorithms, and five prompt-based auditing methods. By using various auditing algorithms, we evaluate the effectiveness and robustness of different unlearning strategies. To explore alternatives beyond prompt-based auditing, we propose a novel technique that leverages intermediate activation perturbations, addressing the limitations of auditing methods that rely solely on model inputs and outputs.

机器遗忘大模型安全隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。