arXiv:2506.17114cs.AI2025-06被引 16

用数学证明检测大模型推理缺陷,发现其真实能力远低于宣称表现。

Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models

  • 构建200个数学证明题数据集,以严格逻辑暴露模型弱点。
  • 超80%题目正确率不足20%,单步推理缺乏严谨性保证。
  • 适合关注模型可信推理与形式化训练的研究者。

大型推理模型(如R1、o3)在数学问题求解上表现出色,但其在主流数据集上的高准确率常被纯数值评估和基准泄露所掩盖,难以反映真实推理能力。为此,我们提出将数学证明的严谨性作为诊断工具,揭示隐藏的失败模式。具体地,引入RFMDataset(Reveal Failure Modes),包含200个多样化的数学证明题,并对先进模型在此数据集上的表现进行详尽评估。深入分析发现10类细粒度错误类型,揭示当前大型推理模型的根本局限:1)模型在数学证明任务上表现不佳,部分问题正确率低于20%,甚至无法完成基础证明;2)模型存在多样的推理失败,尤其缺乏单步推理的正确性与严谨性保障;3)推理过程中存在幻觉与不完整现象。研究发现,模型自省不足以解决当前逻辑困境,亟需形式化且细粒度的逻辑训练。

原文摘要 · Abstract (English)

Large reasoning models (e.g., R1, o3) have demonstrated remarkable mathematical problem-solving abilities. However, the high reported accuracy of these advanced models on popular datasets, reliance on purely numerical evaluation and potential benchmark leakage, often masks their true reasoning shortcomings. To address this, we propose leveraging the inherent rigor and methodological complexity of mathematical proofs as a diagnostic tool to expose these hidden failures. Specifically, we introduce the RFMDataset (Reveal Failure Modes), a collection of 200 diverse mathematical proof problems, and thoroughly evaluate advanced models' performance on it. Our in-depth analysis of their failures uncovers 10 fine-grained error types, which shows fundamental limitations in current large reasoning models: 1) large reasoning models grapple profoundly with mathematical proofs, with some generating entirely correct proofs for less than 20% of problems and failing even on basic ones; 2) models exhibit a diverse spectrum of reasoning failures, prominently demonstrating the lack of guarantees for the correctness and rigor of single-step reasoning; and 3) models show hallucination and incompleteness during the reasoning process. Our findings reveal that models' self-reflection is insufficient to resolve the current logical dilemmas, necessitating formalized and fine-grained logical training.

大模型推理数学证明逻辑缺陷模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。