大模型数学推理准确率高,但多数靠不可靠路径,易沉默出错。
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
- 用新指标发现:正确推理中仅18.4%稳定可靠,81.6%依赖不一致计算路径。
- 70亿参数模型比15亿提升0%准确率,规模扩张无增益。
- 近20%推理具类思维链特征,但8.8%结果自信却错误,需警惕沉默失败。
尽管数学推理模型广泛应用于教育、自动辅导和决策支持系统,但其存在根本性计算不稳定性。我们发现,当前顶尖模型(Qwen2.5-Math-7B)在评估集上达到61%准确率,但其中仅18.4%的正确预测源于稳定、可信的推理路径,81.6%则通过计算不一致的路径产生。此外,8.8%的预测为沉默失败——即高置信度但错误的结果。通过新颖的可信赖性度量进行综合分析,揭示:(1) 推理质量与正确性呈弱负相关(r=-0.21, p=0.002),反映的是二分类阈值效应而非单调反向关系;(2) 参数从1.5B扩展至7B(4.7倍)在所测子集(GSM8K的6%)上未带来任何准确率提升,需在完整基准上验证;(3) 隐空间推理采用多样化计算策略,约20%表现出类思维链模式。这些发现表明,基准准确率可能掩盖计算不可靠性,亟需改革评估体系,引入超越单样本指标的稳定性度量。
原文摘要 · Abstract (English)
Mathematical reasoning models are widely deployed in education, automated tutoring, and decision support systems despite exhibiting fundamental computational instabilities. We demonstrate that state-of-the-art models (Qwen2.5-Math-7B) achieve 61% accuracy through a mixture of reliable and unreliable reasoning pathways: 18.4% of correct predictions employ stable, faithful reasoning while 81.6% emerge through computationally inconsistent pathways. Additionally, 8.8% of all predictions are silent failures -- confident yet incorrect outputs. Through comprehensive analysis using novel faithfulness metrics, we reveal: (1) reasoning quality shows weak negative correlation with correctness (r=-0.21, p=0.002), reflecting a binary classification threshold artifact rather than a monotonic inverse relationship; (2) scaling from 1.5B to 7B parameters (4.7x increase) provides zero accuracy benefit on our evaluated subset (6% of GSM8K), requiring validation on the complete benchmark; and (3) latent reasoning employs diverse computational strategies, with ~20% sharing CoT-like patterns. These findings highlight that benchmark accuracy can mask computational unreliability, demanding evaluation reforms measuring stability beyond single-sample metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。