arXiv:2605.26454cs.CL2026-05

针对不同语言功能设计专门的模型去学习方法,提升安全性和可控性。

Model Unlearning Objectives Vary for Distinct Language Functions

论文配图:Model Unlearning Objectives Vary for Distinct Language Functions
图 1 · 摘自论文原文
  • 为危险知识和毒性内容分别设计独立的去学习目标与方法。
  • 在四个7-8B开源模型上实现显著效果,验证了方法的有效性。
  • 适合关注大模型安全、可控生成的研究者与实践者。

大型语言模型(LLMs)在预训练过程中会习得有害属性,包括危险知识和有毒文本生成。正如后训练通过不同目标塑造不同行为,我们认为去学习方法也应针对具体语言功能进行设计。为此,我们研究了两种机制上不同的去学习目标:危险知识去学习与毒性去学习。针对危险知识,提出一种基于余弦的元学习变体RMU;针对毒性,提出基于层间探测方向的多层目标。在四个7-8B规模的开源模型上,该方法基于各自目标实现了优异表现。结果表明,去学习应被视为一类问题,类似于大模型后训练中的多种任务类型。

原文摘要 · Abstract (English)

Large language models (LLMs) learn undesirable properties during pretraining, including dangerous knowledge and toxic text generation. Just as post-training uses different objectives to shape different behaviors, we argue that unlearning methods should be designed for the language function at issue. To study this, we consider two mechanistically distinct unlearning goals, dangerous-knowledge unlearning and toxicity unlearning. For dangerous knowledge, we introduce a cosine-based, meta-learned variant of RMU. For toxicity, we propose a multi-layer objective based on layer-specific probe directions. Across four open-source 7-8B models, our methods achieve strong results, based on distinct training objectives for the two types of unlearning. Overall, our results suggest that unlearning should be studied as a family of problems, analogous to the multiple types of LLM post-training.

大模型安全去学习语言功能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。