arXiv:2506.15076cs.CLcs.LG2025-06被引 2

训练时的表述方式影响大模型删知识的难易程度。

Learning-Time Encoding Shapes Unlearning in LLMs

  • 用改写句式训练能提升删知识效果
  • 从一段文字中删单条信息很困难
  • 适合关注模型可删除性的研究者

随着大语言模型在现实世界中的广泛应用,事后移除特定知识(即‘删知识’)的能力变得至关重要,原因包括隐私合规、纠正过时或有害内容等。以往工作提出了删知识基准和算法,通常假设训练过程和目标模型固定不变。本文通过实验探究训练阶段的知识编码方式对删知识效果的影响。结果发现:(1) 使用改写表述进行训练可提升删知识性能;(2) 从一段文本中删除单一知识项极具挑战性。这些发现表明,训练阶段的知识编码方式可能在实现可靠的事后删知识中起核心作用。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in the real world, the ability to ``unlearn'', or remove specific pieces of knowledge post hoc, has become essential for a variety of reasons ranging from privacy regulations to correcting outdated or harmful content. Prior work has proposed unlearning benchmarks and algorithms, and has typically assumed that the training process and the target model are fixed. In this work, we empirically investigate how learning-time choices in knowledge encoding impact the effectiveness of unlearning factual knowledge. Our experiments reveal two key findings: (1) learning with paraphrased descriptions improves unlearning performance and (2) unlearning individual piece of knowledge from a chunk of text is challenging. Our results suggest that learning-time knowledge encoding may play a central role in enabling reliable post-hoc unlearning.

大模型删知识训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。