arXiv:2605.27773cs.CLcs.AI2026-05

模型推理过程真能说明它为何改主意吗?研究发现多数推理只是表演。

Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict

  • 通过对比200个问题下8个模型的推理,发现决策变化时理由几乎不变
  • 模型自评信心在冷门事实中仍能预测正确答案,但整体相关性极弱
  • 只有GPT-4o的推理与决策真正挂钩,其他模型多为‘伪解释’

当语言模型遇到与其训练知识矛盾的文档时,必须选择相信文档还是自己。以往研究证明该选择取决于事实的知名程度。本文提出内省忠实性,测试了200个问题、8个模型、4种提示条件下链式思维(CoT)是否真实反映这一机制。结果发现,即使决策相反,推理内容仍高度一致:翻转对保留96%的相同回答相似度(d=0.34;ROUGE-L验证,d=0.45)。然而自评信心仅在冷门事实中携带微弱真实信号:信心可预测决策(p<0.001),并追踪个体知识水平(r=0.134)。GPT-4o是唯一推理与决策统计显著耦合的模型。Claude Sonnet 4.6信心范围最广(SD=1.39),但总体相关性趋近于零,因信心-决策关系在不同提示条件下反转;温度消融实验确认此为模型特异性现象。内部思考标记比对外展示的CoT更具决策敏感性(p=0.033)。CoT由约96%的决策无关知识呈现和一层薄弱但真实的信心层构成。监控时应关注信心值,而非推理过程。

原文摘要 · Abstract (English)

When a language model sees a document contradicting its training knowledge, it must choose: follow the document or trust itself. Prior work proved this choice depends on how well-known the fact is. We ask: does the model's chain-of-thought (CoT) reasoning faithfully report this mechanism? We introduce introspective faithfulness and test it across 200 questions, 8 models, and 4 prompt conditions. We find CoT reasoning is highly stable across opposite decisions: flip pairs retain 96% of same-answer similarity (d=0.34; confirmed by ROUGE-L, d=0.45). Yet self-rated confidence carries a faint genuine signal: for obscure facts where entity fame is uninformative, confidence still predicts decisions (p<0.001) and tracks item-level knowledge (r=0.134). GPT-4o is the only model with statistically reliable reasoning-decision coupling. Claude Sonnet 4.6 shows the widest confidence range (SD=1.39) but near-zero pooled correlation because the confidence-decision relationship reverses between conditions; a temperature ablation confirms this is model-specific. Internal thinking tokens show greater decision-sensitivity than user-facing CoT (p=0.033). CoT decomposes into a decision-invariant knowledge display (~96%) and a thin confidence layer with weak but real signal. For monitoring: read confidence, not the argument.

可解释性大模型推理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。