大模型学新知识时,遇矛盾会破坏80%旧知识,像人一样拒绝冲突才能避免灾难。
In Praise of Stubbornness: An Empirical Case for Cognitive-Dissonance Aware Continual Update of Knowledge in LLMs
- 区分常用与不常用神经元,尝试选择性更新知识
- 学习10-100条矛盾信息就损毁80%无关知识
- 可用简单特征识别矛盾,适合改进模型抗干扰能力
通过系统性实证研究,我们发现大型语言模型存在一个根本性问题:当学习与已有知识不冲突的事实时能保持安全,但若引入矛盾信息,会导致无关知识的灾难性崩溃。与人类自然抗拒矛盾不同,这些模型无差别接受矛盾,造成严重干扰,即使仅更新10-100条矛盾事实,也能破坏高达80%的原有知识。为探究是否可通过选择性可塑性缓解干扰,我们测试了针对高频使用(顽固)和低频使用(灵活)神经元的靶向更新策略。结果发现,对非矛盾更新而言,保留高频神经元可显著提升知识保留率(98%对比标准更新的93%),但对矛盾更新,无论策略如何均引发灾难性干扰。该现象在从GPT-2到GPT-J-6B的多种模型规模中持续存在,表明神经网络处理矛盾存在根本局限。最后,我们证明仅用简单模型特征即可实现95%以上的矛盾信息检测准确率,为构建类似人类、天然抵抗矛盾的新架构提供可能。
原文摘要 · Abstract (English)
Through systematic empirical investigation, we uncover a fundamental and concerning property of Large Language Models: while they can safely learn facts that don't contradict their knowledge, attempting to update facts with contradictory information triggers catastrophic corruption of unrelated knowledge. Unlike humans, who naturally resist contradictory information, these models indiscriminately accept contradictions, leading to devastating interference, destroying up to 80% of unrelated knowledge even when learning as few as 10-100 contradicting facts. To understand whether this interference could be mitigated through selective plasticity, we experiment with targeted network updates, distinguishing between previously used (stubborn) and rarely used (plastic) neurons. We uncover another asymmetry: while sparing frequently-used neurons significantly improves retention of existing knowledge for non-contradictory updates (98% vs 93% with standard updates), contradictory updates trigger catastrophic interference regardless of targeting strategy. This effect which persists across tested model scales (GPT-2 to GPT-J-6B), suggests a fundamental limitation in how neural networks handle contradictions. Finally, we demonstrate that contradictory information can be reliably detected (95%+ accuracy) using simple model features, offering a potential protective mechanism. These findings motivate new architectures that can, like humans, naturally resist contradictions rather than allowing destructive overwrites.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。