arXiv:2609.04290cs.CLcs.AI2026-09

揭示大模型如何整合外部证据,发现其决策受自身偏好影响

Evidence Integration in Large Language Models

  • 证据通过调整初始答案分布影响模型决策,依赖接收者先验与证据倾向
  • 相同证据可能提升弱模型却损害强模型,且自身错误比外部错误更易被接受
  • 证据整合是模型内部特定控制机制,适用于科学推理等复杂任务

尽管大语言模型越来越多地依赖工具、检索增强生成、其他智能体和用户提供的外部证据进行推理,但其如何将这些证据融入已开始形成的决策仍不清晰。本文提出一种分布理论,指出证据会改变接收者对初始答案的分布,由接收者先验权重和候选证据倾斜驱动,得出三项预测:第一,接收者更倾向于采纳自身概率更高的候选答案;第二,更易接受自身特征性错误而非来自其他来源的错误;第三,相同证据可能提升弱模型表现,反而损害强模型。我们在超千万次试验、十二个来自四个家族的大模型及八个领域(包括量子力学、物理、遗传学和分子生物学等四类科学发现任务)中验证了这些预测。该理论还引出接收者相对可靠性边界:与接收者一致的错误导致性能下降更剧烈,即使错误率相同。此外,模型在内部验证后仍能整合无效候选答案(命题约束下93-100%;物理与生命科学推理任务中达99.4%),表明证据整合是基于接收者属性的特定控制策略,而非对证据源的简单信任。因果干预显示,候选整合发生在网络后期,表现为一系列结构化步骤:接纳外部候选、促进其地位、传递至答案状态。验证表示可解码但对答案影响甚微。J-透镜分解表明,口头验证背后的态与候选整合背后的态完全分离。

原文摘要 · Abstract (English)

Despite increasing reliance on LLMs that reason with external evidence supplied by tools, retrieval-augmented generation, other agents, and users, how LLMs integrate such evidence into decisions they have already begun to form remains largely unclear. We present a distributional theory in which evidence shifts the receiver's distribution of initial answers, driven by a receiver prior weight and a candidate evidence tilt, leading to three predictions. First, candidates more probable to the receiver are more persuasive. Second, receivers more readily integrate characteristic errors of their own than foreign errors from different sources. Third, identical evidence can improve weaker models and harm stronger ones. We confirm these over ten million trials, twelve LLMs from four families, and eight domains, four of them scientific discovery tasks in the physical and life sciences: quantum mechanics, physics, genetics, and molecular biology. The law also yields a receiver-relative reliability frontier: receiver-congruent errors depress performance more steeply than random errors of the same rate. LLMs also integrate candidates even after internally verifying their invalidity (93-100% with propositional constraints; up to 99.4% on held-out physical and life-sciences reasoning), demonstrating evidence integration is a receiver-specific control policy over existing distributions, determined by receiver properties rather than scalar trust in the evidence source. Causal interventions show candidate integration is implemented late in the network, as a structured sequence of steps admitting external candidate answers, promoting them, and transporting them into the answer state. Representations of verification are decodable but have little causal impact on answers. A J-lens decomposition shows the state underlying verbalized verification is fully dissociable from that underlying candidate integration.

大模型推理证据整合认知机制科学发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。