arXiv:2601.09445cs.CLcs.AI2026-01被引 5

揭示大模型内部知识冲突的定位与解决机制

Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models

  • 构建框架定位模型内部知识冲突位置
  • 真实冲突比合成冲突更难通过干预解决
  • 冲突由独立电路编码,无统一处理路径

在语言模型中,当同一主题的不一致信息被编码进参数化知识时,会产生内部知识冲突。以往研究主要关注模型内部知识与外部来源之间的冲突(上下文-记忆冲突),而对内部冲突的理解仍不充分。本文设计了一个框架,用于识别语言模型中内部冲突的编码位置。在四个模型上使用合成与真实世界知识冲突进行测试,发现冲突通常在所有模型的最后几层出现并被解决,但针对真实冲突的干预效果显著较差。定向注意力头干预优于逐层干预,过滤分析显示,针对单一竞争事实的专用头在合成冲突中更为常见,解释了这一差距。此外,未发现处理冲突的通用电路,结果表明不同竞争知识可能由独立电路分别编码,从而引发冲突。本研究首次提供了内部知识冲突机制的解析,凸显了合成与真实场景间的巨大差距。

原文摘要 · Abstract (English)

In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the model's parametric knowledge. Prior work has primarily focused on resolving conflicts between a model's internal knowledge and external sources, which is known as context-memory knowledge conflict, through approaches such as fine-tuning or knowledge editing, while the understanding of conflicts that arise internally remains largely unexplored. In this work, we design a framework to identify where internal conflicting knowledge is encoded within LMs. We test our framework on four LMs using both synthetic and real-world knowledge conflicts. We find that internal conflicts often arise and are resolved in the final layers across all models, but that interventions are markedly less effective on real-world knowledge conflicts. Targeted attention-head interventions outperform layer-wise ones, and a filtering analysis shows that heads specialized for a single competing fact are far more common in synthetic conflicts, helping explain this gap. Finally, we find no evidence of a single universal circuit for handling knowledge conflict. Instead, our results suggest that distinct circuits may separately encode competing pieces of knowledge, giving rise to conflict. Our results offer a first mechanistic account of intra-memory conflict resolution and highlight a substantial gap between synthetic and real-world settings.

大模型知识冲突机制研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。