arXiv:2605.27083cs.CLcs.CR2026-05

发现大模型反事实训练存在隐藏代价,导致幻觉增加和优化困难。

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

论文配图:On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
图 1 · 摘自论文原文
  • 通过反事实数据训练模型生成虚构知识以删除旧信息
  • 发现反事实训练引发知识冲突与幻觉溢出,导致幻觉率上升
  • 提出新评测框架RWKU+,可诊断训练中的梯度问题

反事实微调(CFT)作为一种大语言模型(LLM)遗忘的新范式,通过训练模型在不想要的内容上生成替代的虚构知识来实现遗忘。然而,本文发现该范式在某些方面仍逊于其他方法,并揭示了两个此前被忽视的隐患:(1)知识冲突——反事实语料内部的不一致引发矛盾梯度,干扰参数优化;(2)幻觉溢出——拟合虚假目标会引入持续存在的编造偏见,导致无关领域幻觉率上升。为系统诊断这些问题,我们提出了扩展基准RWKU+,包含新型权衡指标与梯度级诊断工具。本文还讨论了该范式的局限性与开销,旨在为更严谨的LLM遗忘研究提供洞见与实用指导。

原文摘要 · Abstract (English)

Counterfactual tuning (CFT) has emerged as a promising paradigm for Large Language Model (LLM) unlearning by training models to generate alternative fictitious knowledge in place of undesired content. However, in this work, we find that this paradigm still underperforms other paradigms in some aspects, and identify two previously overlooked pitfalls underlying this gap: (1) knowledge conflict, where mutual inconsistencies within counterfactual corpora induce conflicting gradients that disrupt parameter optimization, and (2) hallucination spillover, where fitting false targets instills a persistent fabrication bias, inflating hallucination rates on unrelated domains. To systematically diagnose these issues, we introduce RWKU+, an extended benchmark equipped with novel trade-off metrics and gradient-level diagnostic tools. Our work further discusses the limitations and overhead of the paradigm, aiming to provide insights and actionable guidance for more rigorous LLM unlearning research.

大模型遗忘反事实训练幻觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。