arXiv:2410.21750cs.CLcs.AI2024-10被引 5

研究大模型如何记住与遗忘虚假知识,发现冲突常识的知识最持久。

Learning and Unlearning of Fabricated Knowledge in Language Models

  • 通过新数据集注入事实,测试不同知识类型的记忆稳定性。
  • 违背常识的事实可维持数万步训练,普通或随机事实迅速遗忘。
  • 冲突知识会引发非目标幻觉,但可通过稀疏更新有效清除。

当新知识被注入大语言模型的训练数据后,其记忆持续多久?我们通过一个新探测数据集「Outlandish」来研究这一问题,该数据集支持测试多种类型的事实。结果显示,在事实新颖性介于符合常识与完全随机之间的某个区间时,记忆最为持久:违背常识的事实可维持数万次训练步骤,而普通事实和随机打乱的事实则快速遗忘。此外,违背常识的事实会“引导”模型在无关提示下产生幻觉,表现出强非目标泛化能力,而普通与随机事实的引导作用较弱。最后,尽管冲突知识影响持久,但通过多步稀疏更新即可大幅清除,且不影响模型训练能力。该方法对缓解数据投毒攻击具有直接应用价值。

原文摘要 · Abstract (English)

What happens when a new piece of knowledge is introduced into the training data and how long does it last while a large language model (LM) continues to train? We investigate this question by injecting facts into LMs from a new probing dataset, "Outlandish", which is designed to permit the testing of a spectrum of different fact types. When studying how robust these memories are, there appears to be a sweet spot in the spectrum of fact novelty between consistency with world knowledge and total randomness, where the injected memory is the most enduring. Specifically we show that facts that conflict with common knowledge are remembered for tens of thousands of training steps, while prompts not conflicting with common knowledge (mundane), as well as scrambled prompts (randomly jumbled) are both forgotten much more rapidly. Further, knowledge-conflicting facts can "prime'' how the language model hallucinates on logically unrelated prompts, showing their propensity for non-target generalization, while both mundane and randomly jumbled facts prime significantly less. Finally, we show that impacts of knowledge-conflicting facts in LMs, though they can be long lasting, can be largely erased by novel application of multi-step sparse updates, even while the training ability of the model is preserved. As such, this very simple procedure has direct implications for mitigating the effects of data poisoning in training.

大模型记忆知识遗忘数据投毒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。