提出三类冲突处理机制,解释大模型为何有时听训练、有时听上下文。
Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation
- 区分三种不同情境:更新、竞争、任务选择,各由不同因素主导。
- 实验证明模型在矛盾时96%听上下文,但取决于具体情境与任务要求。
- 框架可解释模型行为差异,适合研究大模型推理机制的学者使用。
大语言模型在处理训练知识与新文档冲突时表现出矛盾行为:部分研究发现模型近一半时间固执于训练答案,另一些研究则显示模型96%时间遵循文档。我们指出,这些矛盾源于未区分三种不同处理情境。本文提出三阶段框架:第一阶段(单源更新,主导因素为证据一致性)、第二阶段(竞争整合,主导因素为参数确定性)、第三阶段(任务适配选择,主导因素为任务知识需求)。通过区分参数强度(暴露频率)与参数唯一性(编码一致性),实证发现二者几乎无关(r = -0.002, p = .97),且强度在稳定事实领域起决定作用。在五种模型(Claude Sonnet 4.6、GPT-5.5、Gemini 2.5 Flash、Llama 4 Maverick、DeepSeek V3)上进行9,970次API调用,三阶段实验验证框架有效性。广义线性回归确认所有模型在第二阶段均存在显著确定性梯度(beta = -0.38 至 -0.50,p ≤ .013,BH-FDR校正)。第三阶段消融实验显示,仅任务表述变化即可使上下文遵循率从接近100%(上下文知识条件)降至6%-71%(参数知识条件),五模型均显著(p < .001)。确定性梯度在多分类建模、抗模糊响应敏感性分析及FDR校正下依然稳健。
原文摘要 · Abstract (English)
The literature on how large language models handle conflict between their training knowledge and a contradicting document presents a persistent empirical contradiction: some studies find models stubbornly retain their trained answers, ignoring provided documents nearly half the time, while others find models readily defer to the document, following context approximately 96% of the time. We argue these contradictions dissolve once one recognises that prior experiments have studied three qualitatively distinct processing situations without distinguishing them. We propose a three-regime framework: Regime 1 (single-source updating, dominant predictor: evidence coherence), Regime 2 (competitive integration, dominant predictor: parametric certainty), and Regime 3 (task-appropriate selection, dominant predictor: task knowledge requirement). We formalise a distinction between parametric strength (exposure frequency) and parametric uniqueness (encoding consistency), showing empirically that these are orthogonal dimensions (r = -0.002, p = .97) with strength as the operative predictor in stable factual domains. We validate the framework across Claude Sonnet 4.6, GPT-5.5, Gemini 2.5 Flash, Llama 4 Maverick, and DeepSeek V3 using 9,970 API calls in three experimental phases. GEE logistic regression confirms the predicted Regime 2 certainty gradient for all five models (beta = -0.38 to -0.50, all p <= .013, BH-FDR corrected). A Regime 3 ablation shows task framing alone flips context-following from near-100% (contextual knowledge condition) to 6-71% (parametric knowledge condition), with all five models significant (p < .001). The certainty gradient is robust to multinomial outcome modeling, sensitivity analyses for hedging responses, and FDR correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。