arXiv:2510.07626cs.LGcs.CL2025-10被引 2

系统梳理大模型遗忘技术,揭示评估漏洞并提出更真实测试方法。

LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics

  • 按方法原理将12种遗忘技术分为三类:差异优化、表征错位、拒绝式遗忘。
  • 发现现有评测依赖选择题准确率,易高估效果,真实生成行为常被忽视。
  • 提出开放问答新指标,揭示不同方法在遗忘与性能间的权衡关系。

大型语言模型(LLMs)的机器遗忘旨在移除不希望的数据、知识和行为(如安全、隐私或版权问题),同时保留有用模型能力。尽管过去两年进展迅速,但该领域研究仍碎片化,对有效遗忘的定义及严谨评估缺乏共识。本文提出一个包含十二种近期有状态遗忘方法的系统性分类,分为三类方法族:基于差异的优化、表征错位和基于拒绝的目标遗忘。基于此分类,我们重新审视了遗忘有效性(UE)、能力保留度(UT)和鲁棒性(Rob)的评估,聚焦于WMDP基准。分析显示,当前评估主要依赖多选题(MCQ)准确率,视角狭窄,常夸大成功,忽略模型实际生成行为。为此,我们引入开放问答(Open-QA)评估指标,更真实反映生成表现,并揭示了不同方法族间的固有UE-UT权衡。此外,我们表明鲁棒性需细粒度分析:例如,在域内重训练与域外微调攻击中,模型脆弱性差异显著,尽管两者均属模型级攻击。本研究旨在为大模型遗忘提供全栈重审与未来方法设计的可操作指导。

原文摘要 · Abstract (English)

Machine unlearning for large language models (LLMs) aims to remove undesired data, knowledge, and behaviors (e.g., for safety, privacy, or copyright) while preserving useful model capabilities. Despite rapid progress over the past two years, research in LLM unlearning remains fragmented, with limited clarity on what constitutes effective unlearning and how it should be rigorously evaluated. In this work, we present a principled taxonomy of twelve recent stateful unlearning methods, grouped into three methodological families: divergence-driven optimization, representation misalignment, and rejection-based targeted unlearning. Building on this taxonomy, we revisit the evaluation of unlearning effectiveness (UE), utility retention (UT), and robustness (Rob), focusing on the WMDP benchmark. Our analysis shows that current evaluations, dominated by multiple-choice question (MCQ) accuracy, offer only a narrow perspective, often overstating success while overlooking the model's actual generation behavior. To address this gap, we introduce open question-answering (Open-QA) metrics that better capture generative performance and reveal the inherent UE-UT tradeoff across method families. Furthermore, we demonstrate that robustness requires finer-grained analysis: for example, vulnerabilities differ substantially between in-domain relearning and out-of-domain fine-tuning, even though both fall under model-level attacks. Through this study, we hope to deliver a full-stack revisit of LLM unlearning and actionable guidance for designing and evaluating future methods.

大模型遗忘学习评估指标鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。