arXiv:2609.08067cs.CL2026-09

热门知识在模型更新时更易出错且传播错误。

Popular Knowledge Propagates More Errors in LLM Knowledge Updating

论文配图:Popular Knowledge Propagates More Errors in LLM Knowledge Updating
图 1 · 摘自论文原文
  • 发现高连接实体相关事实在更新中更易被污染。
  • 热门事实错误会通过知识图谱扩散至更多关联事实。
  • 提出轻量级策略保留热门事实,减少遗忘。

通过微调更新语言模型的知识对保持输出时效性至关重要,但也可能引发事实遗忘和新幻觉。已有研究指出长尾知识难以获取且难以留存。本文研究一个互补问题:在模型已正确编码的事实中,哪些最易在其他更新中受到连带污染?为在真实事实分布下探究此问题,我们构建了大规模图 FACTPROP,通过共享头实体或尾实体链接三元组,保留知识间的关联。在事实语句上微调模型,并测量每次更新后的正确转错误事实。结果揭示与以往长尾脆弱性不同的模式:已正确编码的事实中,与高度连接实体相关的事实更易被邻近更新所腐蚀,且此类错误传播范围更广。结构流行度可预测脆弱性与下游损害。受此启发,我们提出轻量级重演策略 Popularity-based Anchoring(PopAnchor),通过保留少量热门事实来减少遗忘。

原文摘要 · Abstract (English)

Updating a language model's knowledge through fine-tuning is essential for keeping its outputs current, yet can also induce factual forgetting and new hallucinations. Prior work shows that long-tail knowledge is harder to acquire and newly memorized long-tail facts are difficult to retain during later fine-tuning. We study a complementary question: among facts that a model has encoded correctly, which are most vulnerable to collateral corruption during other updates? To investigate this question under a realistic factual distribution, we construct a large-scale graph FACTPROP of verified Wikipedia facts by linking triples that share head or tail entities, thereby preserving connections among factual knowledge. We fine-tune models on factual statements and measure correct-to-incorrect facts after each update. Our results reveal a pattern distinct from prior findings on long-tail vulnerability during acquisition and retention: among facts that models already answer correctly, those associated with highly connected entities are more likely to be corrupted by neighboring updates, and updates to such facts propagate errors more broadly. Structural popularity therefore predicts both vulnerability and downstream damage. Inspired by this finding, we propose Popularity-based Anchoring (PopAnchor), a lightweight rehearsal strategy that preserves a small set of popular facts and reduces forgetting.

知识更新幻觉传播模型遗忘知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。