发现知识纠缠度决定大模型删不掉的错误信息,可提前预测并干预。
What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs
- 用知识纠缠度衡量事实与模型其他知识的关联强度。
- 不同删除算法下,纠缠度与泄露相关性从正转负,甚至反转。
- 能通过操控纠缠度验证因果关系,实现未删前预判泄露风险。
大语言模型的去学习通常以是否能恢复被删除的事实来评估。本文关注的是:一个事实与其知识结构的纠缠程度,能否预测其在去学习后是否泄露?这种关系是否系统性变化?是否具有因果性?在多种去学习算法(WHP 和 GA+KL)、两个领域(虚构的哈利·波特与非虚构的2000–2010年美国参议员)及四个模型(27亿至130亿参数)中,我们发现:去学习前,纠缠度越高的事实被回忆的概率越高(相关系数 r = +0.39 至 +0.51)。WHP 算法削弱但仍保持正相关(r = +0.16 至 +0.33);而 GA+KL 在所有场景中均使相关性反转(r = -0.14 至 -0.25),这是目前文献中首次报告此类反转。为验证因果性,我们直接操纵提示的纠缠得分,固定内容和模型,结果显示回忆行为按预期方向移动,并在 GA+KL 下发生符号反转,与相关分析一致。此操作证明去学习作用于底层知识结构,而非仅输出结果。基于此,我们训练了一个预测模型,可在去学习前估算提示的后验事实性/幻觉特征,为模型审计提供一种新型提示风险筛查方法。
原文摘要 · Abstract (English)
Unlearning in large language models (LLMs) is usually evaluated as whether an "unlearned" fact can be recovered. We instead ask whether a fact's structural entanglement with the rest of a model's knowledge predicts whether it leaks after unlearning, whether this relationship changes systematically, and whether it is causal. Across varied unlearning algorithms (WHP and GA+KL), two domains (both fictional Harry Potter and non-fictional U.S. Senators, 2000-2010), and four models (2.7B-13B parameters), we find that before unlearning, more entangled facts are recalled more often (r = +0.39 to +0.51). WHP weakens this relationship but stays positive (r = +0.16 to +0.33); GA+KL inverts it in every domain and model size (r = -0.14 to -0.25). To our knowledge, this is the first report of this specific reversal in the unlearning literature. To confirm this is causal beyond correlation, we directly manipulate a prompt's entanglement score, holding its content, target model fixed, and show recall moves in the predicted direction and reverses sign under GA+KL in the same direction as the correlational analysis. This manipulation is our central evidence that unlearning acts on the underlying knowledge structure, not just its output. Building on this, we train a predictive model that estimates a prompt's post-unlearning factuality/hallucination profile before unlearning is run, creating a unique way of model auditing in triaging which prompts are likely to leak.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。