大模型解题强但评理弱,易被错误推理骗过。
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

- 用瑕疵推理+正确答案的测试集,专测模型评理能力。
- 顶尖模型评理准确率仅48%,远低于解题近100%正确率。
- 模型偏信答案对错,忽略推理过程,适合研究模型可信度。
人类在评估推理能力上通常优于原创推理,而大模型(LRM)则以生成复杂推理链见长。我们通过VAIR数据集——包含推理有误但答案正确的数学题——考察大模型的推理评价能力,以排除生产干扰。结果发现,人类在评分上仅比解题差6%,而大模型存在显著的生产-评价差距:前沿模型在评估VAIR解题时准确率低至48%,尽管其解题正确率接近100%。通过思维链分析,我们发现模型存在答案确认偏差:它们倾向于先产出答案,再检查是否匹配,而非逐步验证推理;即使察觉异常,也会编造合理化解释。线性探测显示,模型激活虽能编码部分有效推理特征,却无法稳健识别无效解法。因果修补最终答案表示后,模型判断与激活状态发生反转,证实答案正确性是导致偏差的主因。这揭示了主流推理训练方法的根本缺陷:激励模型生成并确认答案,却未强化对推理过程本身的严谨评估。
原文摘要 · Abstract (English)
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trained to excel at producing long chains of reasoning to solve complex problems. How then do LRMs perform at evaluating reasons? We investigate this with the Valid-Answer-Invalid-Reasoning (VAIR) dataset: math problems and solutions with trivial reasoning flaws but valid answers, designed to isolate reasoning evaluation from the confound of reasoning production. Unlike humans, who we find are only 6% worse at grading than solving such problems, we find a substantial production-evaluation gap in LRMs: frontier models score as low as 48% when evaluating VAIR solutions, despite near-perfect solution production. Why this enigma? Through chain-of-thought (CoT) analysis, we find evidence of an answer confirmation bias: LRMs often produce then check for the correct answer instead of carefully verifying each step, fabricating rationalizations even when noticing anomalous reasoning. Linear probes corroborate this, showing that while LRM activations encode some representation of valid reasoning, they fail to robustly represent VAIR solutions as invalid. Causal patching of the final answer's representations causes LRM verdicts and activations to flip, demonstrating that answer validity is responsible for models' confirmation biases. These findings indicate an outstanding limitation in dominant approaches to reasoning training, which incentivize LRMs to produce and confirm reasoning towards correct answers, but not to robustly evaluate the underlying reasons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。