arXiv:2606.27379cs.CLcs.AI2026-06被引 1

警告:大模型领域滥用‘机器遗忘’术语,应明确区分真实数据删除与其他安全策略。

Position: The Term "Machine Unlearning" Is Overused in LLMs

  • 真实遗忘指移除特定训练数据影响,使模型表现近似于重新训练
  • 当前许多‘遗忘’任务实为对齐、抑制或编辑,目标不同需不同命名
  • 术语混淆导致评估标准错用,表面不泄露未必真实现等效重训

大型语言模型面临日益增长的删除需求,包括合规性要求、版权争议及安全政策。本文指出,'机器遗忘'一词在大模型研究中被过度泛化,应仅用于定义明确的数据集删除:移除特定遗忘数据的影响,使模型结果近似于未包含该数据的重新训练。我们主张,当前被标记为'遗忘'的诸多任务(如拒绝有害请求、知识实体移除、定向压制)实则追求不同目标,多依赖政策而非技术等价,需采用对齐、抑制、编辑或混淆等不同术语与基线。这种混淆非表面问题:因论文对同一标签隐含不同保证,现有指标和基准常被误用于不匹配场景,仅奖励表层不披露(如低ROUGE/遗忘准确率),却未验证重训等价性或能力残留情况。文章呼吁建立更严格的技术术语体系,绑定明确承诺与参考模型,并推动与声称目标一致的评估方式。

原文摘要 · Abstract (English)

Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/licensing disputes, and safety or product-policy requirements. This position paper argues that machine unlearning is overused as a term in LLM research and should be reserved for dataset-defined deletion: removing the training influence of a precisely specified forget set such that the resulting model is approximately indistinguishable from retraining without that data. We contend that many tasks currently labeled "unlearning" (e.g., refusal for harmful requests, entity/knowledge removal, or targeted suppression) pursue different, often policy-dependent objectives and therefore require different terminology and baselines (e.g., alignment, suppression, editing, obfuscation). We further argue that this confusion is not cosmetic: because papers make different implicit guarantees under the same label, metrics and benchmarks are frequently reused outside their intended scope, rewarding surface-level non-disclosure (e.g., low ROUGE/forget accuracy) even when retraining-equivalence is not tested and derived capabilities remain. We conclude by calling for stricter terminology tied to explicit guarantees and reference models, and for evaluations that match the claimed objective.

机器遗忘术语规范大模型安全评价基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。