arXiv:2609.03148cs.CL2026-09

研究大模型如何解决上下文内部知识冲突,提出新数据集与改进方法。

Large Language Models in Resolving Contextual Knowledge Conflicts

论文配图:Large Language Models in Resolving Contextual Knowledge Conflicts
图 1 · 摘自论文原文
  • 构建六类上下文冲突分类体系,建立包含5781样本的数据集。
  • 发现当前模型在处理冲突时仍表现不佳,且偏好早期信息。
  • 提出无需标签的激活引导方法,提升推理与摘要质量。

以往研究多关注大模型内部参数知识与外部上下文之间的冲突,而本文聚焦于上下文自身知识的内在冲突。我们提出了六类上下文冲突类型(事实、推断、时间、粒度、视角、模糊性),并构建了涵盖5781个样本的综合数据集ContextConflict,覆盖推理与摘要任务,包含显式矛盾和需多步推理的隐式冲突。在九个大模型上的实验表明,现有模型在解决此类冲突上仍有明显不足。通过机制可解释分析,揭示了模型对冲突的潜在感知及背后的表征几何结构。此外,发现模型存在一致的早期证据偏好,这是有效冲突解决的主要障碍。基于此,我们提出一种无需训练、无需标签的激活引导方法,可促进更全面地融合证据,在本数据集上显著提升推理准确率,并生成更高质量、更平衡的摘要。

原文摘要 · Abstract (English)

Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce a taxonomy of six types of contextual conflicts (factual, inferential, temporal, granularity, perspective, and ambiguity) and contribute a comprehensive dataset ContextConflict for this setting. The dataset contains 5,781 samples, covers both reasoning and summarization tasks, and includes both explicit contradictions and implicit conflicts that require multi-step reasoning. Experiments on nine LLMs show that current models still fall short in resolving contextual knowledge conflicts. We further provide mechanistic interpretability insights into how LLMs process such conflicts, revealing their latent awareness of conflicts and the representational geometry underlying conflict processing. In addition, our analysis uncovers a consistent model bias towards earlier evidence, and this positional preference serves as a key obstacle to effective conflict resolution. Motivated by these findings, we further propose a simple training-free, label-free steering method that steers activations to encourage a more comprehensive incorporation of evidences for better conflict resolution. On our dataset, the method consistently improves accuracy on reasoning tasks and generates higher-quality, more balanced summaries for summarization tasks.

大模型知识冲突可解释性推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。