arXiv:2601.21702cs.LGcs.CL2026-01被引 1

通过重定向记忆表征,让模型遗忘的同时还能可控地控制行为和能力。

Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities

  • 用目标向量重定向遗忘样本的隐藏表征,实现可控遗忘。
  • 实验验证了遗忘后模型在真伪、情感、拒绝回答等方面可被精确控制。
  • 该技术既可能带来风险,也可用于提升模型的上下文学习能力。

我们研究了表示误导(Representation Misdirection, RM)这一类大语言模型遗忘方法,其通过将遗忘样本的潜在表征引导至目标向量来实现遗忘。尽管重要,但目标向量的作用仍不明确。本文基于线性表征假设,提出若能识别对应高层概念的一维表征,则可在遗忘表征空间中对这一概念向量进行线性操作。在此视角下,我们假设:超越遗忘本身,通过RM实现的机器遗忘还会诱发与高层概念相关的可控涌现行为和更强的辅助能力。实证结果在多种任务中验证了该假设,包括行为控制(如控制模型的真伪性、情感倾向、拒答意愿和语言风格)和能力增强(如提升模型的上下文学习能力)。研究发现,该现象若被滥用可能构成隐性风险,但也可被有意识利用以开发具备更强能力和可控行为的遗忘模型。

原文摘要 · Abstract (English)

We consider Representation Misdirection (RM), a class of large language model (LLM) unlearning methods that achieve forgetting by redirecting the forget-representations, that is, latent representations of forget-samples, toward a target vector. Despite being important, the roles of the target vector used in RM, however, remain underexplored. Here, we approach and revisit RM through the lens of the Linear Representation Hypothesis. Specifically, if one can identify a one-dimensional representation corresponding to a high-level concept, the Linear Representation Hypothesis enables linear operations on this concept vector within the forget-representation space. Under this view, we hypothesize that, beyond forgetting, machine unlearning via RM elicits controllable emergent side behaviors and stronger side capabilities corresponding to the high-level concept. Our hypothesis is empirically validated across a wide range of tasks, including behavioral control (e.g., controlling unlearned models' truthfulness, sentiment, refusal, and language) and capability enhancement (e.g., improving unlearned models' in-context learning (ICL) capability). Our findings reveal that this phenomenon could be either a hidden risk if misused or a mechanism that can be harnessed for developing unlearned models that require stronger capabilities and controllable behaviors.

机器遗忘可控行为大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。