arXiv:2606.05403cs.LGcs.AI2026-06

大模型能识破假数据,却在综合多源信息时无视真假,只看表达方式是否像真。

Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation

论文配图:Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation
图 1 · 摘自论文原文
  • 模型识别单个假数据有效,但合成多源信息时忽略数值真实性
  • 无论数据真假,只要表达风格像分析文本就给予相同权重
  • 提示词改进无效,反而导致全盘怀疑,无法精准分辨真伪

语言模型日益充当认知代理,从多个来源综合证据以支持决策。然而,它们是评估证据质量,还是仅根据表面呈现进行聚合,仍不明确。我们发现,模型具备独立识别伪造统计数据的能力,但在多源信息合成过程中并未调用该能力,无论统计数字是否伪造,均生成相似的数值估计。具体而言,源影响力由‘方法论-语域’门控机制决定,该机制响应分析性文本的分布特征,却不关注数值有效性:例如,统计上不可能的置信区间与有效区间获得相同权重。这种行为分离在六个来自四个模型家族(Anthropic Claude、Qwen、OLMo、OpenAI GPT-5.4)和三个专业领域中均复现。机制分析(包括因果追踪、线性探测和组件归因)一致表明,模型编码并因果使用‘方法论-语域’表示,可跨领域迁移;而数值有效性信号虽可在孤立状态下解码,但在多源合成中被抑制至随机水平。基于提示的缓解策略,即使使用包含确切统计检查的黄金清单,也仅产生全面怀疑而非选择性辨别,且所考察的后训练流水线强化了此捷径,未建立数值验证能力。此失败不源于迎合用户偏好,而是追踪来源是否呈现为分析可信,而非其主张是否自洽。我们称之为‘认识论对齐’:与偏好对齐和安全对齐类似,问题不在能力,而在部署。

原文摘要 · Abstract (English)

Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evaluate the quality of that evidence, or merely aggregate it based on surface presentation, remains poorly understood. We show that models possess the capability to detect fabricated statistics in isolation but do not recruit this capability during multi-source synthesis, producing similar numeric estimates whether the statistics are fabricated or valid. Specifically, source influence is governed by a methodology-register gate that responds to the distributional register of analytical text but not to numeric validity: for example, statistically impossible confidence intervals receive the same weight as valid ones. The behavioral dissociation replicates across six models from four families (Anthropic Claude, Qwen, OLMo, and OpenAI GPT-5.4) and three professional domains. Mechanistic analyses, including causal tracing, linear probes, and component-level attribution, converge on the same account: the model encodes and causally uses a methodology-register representation that transfers across domains, while numeric-validity signals, decodable in isolation, are suppressed to chance during multi-source synthesis. Prompting-based mitigations, even an oracle checklist naming the exact statistical checks, produce blanket skepticism rather than selective discernment, and the post-training pipelines we examine reinforce the shortcut without building numeric verification. Unlike sycophancy, which tracks user preference, this failure tracks whether a source presents as analytically credible, not whether its claims are consistent. We term this \textit{epistemic alignment}: like preference and safety alignment, the question is not capability but deployment.

大模型推理认知对齐可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。