arXiv:2510.02370cs.CLcs.AI2025-10ACL被引 3

训练数据如何影响大模型对记忆知识和上下文知识的使用方式

How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models

  • 通过合成数据实验,发现三类看似有害的数据特征协同促进知识平衡使用
  • 高置信度事实依赖参数化知识,陌生信息则依赖上下文,且该机制可被数据设计调控
  • 研究结果适用于真实预训练数据,对模型微调策略有指导意义

大语言模型同时利用预训练中习得的参数化知识与推理时提供的上下文知识。当两者冲突时,模型根据内部置信度进行仲裁:对高置信度事实倾向于使用参数化知识,对不熟悉内容则依赖上下文。然而,促成这一行为的训练条件尚不明确。本文通过使用合成语料库开展受控实验,识别塑造知识利用模式的具体数据属性。结果揭示一个反直觉发现:稳健、均衡地使用两类知识是一种涌现特性,其形成需同时具备三个通常被视为负面的因素:(i) 文档内重复,(ii) 中等程度的文档内不一致,(iii) 知识分布偏斜。我们进一步证明这些动态现象存在于真实世界的大规模语言模型预训练过程中,并分析了后训练过程如何重塑仲裁策略。研究为设计支持参数化与上下文知识可靠融合的训练数据提供了实证依据。

原文摘要 · Abstract (English)

Large language models leverage both parametric knowledge acquired during pretraining and in-context knowledge provided at inference time. Crucially, when these sources conflict, models arbitrate based on their internal confidence, preferring parametric knowledge for high-confidence facts while deferring to context for less familiar ones. However, the training conditions that give rise to these fundamental behaviors remain unclear. Here we conduct controlled experiments using synthetic corpora to identify the specific data properties that shape knowledge utilization. Our results reveal a counterintuitive finding: the robust, balanced use of both knowledge sources is an emergent property that requires the co-occurrence of three factors typically considered detrimental, including (i) intra-document repetition, (ii) a moderate degree of intra-document inconsistency, and (iii) a skewed knowledge distribution. We further show that these dynamics arise in real-world language model pretraining and analyze how post-training procedures reshape arbitration strategies. Together, our findings provide empirical guidance for designing training data that supports the reliable integration of parametric and in-context knowledge in language models.

知识融合训练数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。