arXiv:2506.09227cs.LGcs.CR2025-06被引 10

区分知识删除与行为抑制,为大模型数据擦除提供新思路

SoK: Machine Unlearning for Large Language Models

  • 按删除意图分类:真正移除知识 or 仅压制输出行为
  • 发现多数方法实际是抑制而非彻底清除
  • 适合关注隐私合规与模型可解释性的研究者

大语言模型(LLM)的机器去学习已成为机器学习关键议题,旨在消除特定训练数据或知识的影响,而无需从头重新训练。已有多种技术被提出,包括梯度上升、模型编辑和隐藏表示重导向等。现有综述多按技术特征分类,却忽略了更根本的维度:去学习的意图——是真正移除内部知识,还是仅抑制其行为表现。本文提出一种基于意图导向的新分类体系。基于此,我们做出三项关键贡献:首先,重新审视近期研究指出许多移除方法在功能上更像抑制,并探讨真正移除是否必要或可行;其次,系统调研现有评估策略,识别当前度量标准与基准的局限性,建议发展更可靠且与意图对齐的评估方法;第三,指出可扩展性及支持连续去学习等实际挑战,制约了去学习方法的广泛应用。本文为理解与推进生成式AI中的去学习提供了综合性框架,旨在支持未来研究并指导数据删除与隐私相关的政策制定。

原文摘要 · Abstract (English)

Large language model (LLM) unlearning has become a critical topic in machine learning, aiming to eliminate the influence of specific training data or knowledge without retraining the model from scratch. A variety of techniques have been proposed, including Gradient Ascent, model editing, and re-steering hidden representations. While existing surveys often organize these methods by their technical characteristics, such classifications tend to overlook a more fundamental dimension: the underlying intention of unlearning--whether it seeks to truly remove internal knowledge or merely suppress its behavioral effects. In this SoK paper, we propose a new taxonomy based on this intention-oriented perspective. Building on this taxonomy, we make three key contributions. First, we revisit recent findings suggesting that many removal methods may functionally behave like suppression, and explore whether true removal is necessary or achievable. Second, we survey existing evaluation strategies, identify limitations in current metrics and benchmarks, and suggest directions for developing more reliable and intention-aligned evaluations. Third, we highlight practical challenges--such as scalability and support for sequential unlearning--that currently hinder the broader deployment of unlearning methods. In summary, this work offers a comprehensive framework for understanding and advancing unlearning in generative AI, aiming to support future research and guide policy decisions around data removal and privacy.

大模型数据删除隐私评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。