新基准评估大模型数学推理结构能力,发现答案正确不等于真会推理。
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
- 设计150道需多约束协调与逻辑构建的数学题,强调推理过程而非答案。
- 模型答案准确率最高5.8/10,但过程评分平均仅4.36/10,揭示评测偏差。
- 引入过程级评分机制,适合研究模型真实推理能力的学者使用。
当前大语言模型在多数数学推理基准上已接近饱和准确率,引发对其真实推理能力的担忧。这种饱和主要源于现有数据集以模板化计算和浅层算术分解为主,缺乏对多约束协调、构造性逻辑合成及空间推理等高级能力的覆盖。为此,我们提出 ReasoningMath-Plus,一个包含150道精心设计的问题的基准,每题聚焦于相互作用的约束、构造性解法形成或非平凡的结构洞察,并附有最小推理骨架,支持细粒度的过程评估。同时,我们引入 HCRS(危险感知链式规则评分)这一确定性步骤级评分函数,并基于标注的推理轨迹训练了过程奖励模型(PRM)。实证表明,尽管领先模型最终答案准确率可达5.8/10,但基于HCRS的整体评估得分平均仅为4.36/10,最佳仅5.14/10,说明仅看答案会严重高估模型的推理鲁棒性。
原文摘要 · Abstract (English)
Recent large language models (LLMs) achieve near-saturation accuracy on many established mathematical reasoning benchmarks, raising concerns about their ability to diagnose genuine reasoning competence. This saturation largely stems from the dominance of template-based computation and shallow arithmetic decomposition in existing datasets, which underrepresent reasoning skills such as multi-constraint coordination, constructive logical synthesis, and spatial inference. To address this gap, we introduce ReasoningMath-Plus, a benchmark of 150 carefully curated problems explicitly designed to evaluate structural reasoning. Each problem emphasizes reasoning under interacting constraints, constructive solution formation, or non-trivial structural insight, and is annotated with a minimal reasoning skeleton to support fine-grained process-level evaluation. Alongside the dataset, we introduce HCRS (Hazard-aware Chain-based Rule Score), a deterministic step-level scoring function, and train a Process Reward Model (PRM) on the annotated reasoning traces. Empirically, while leading models attain relatively high final-answer accuracy (up to 5.8/10), HCRS-based holistic evaluation yields substantially lower scores (average 4.36/10, best 5.14/10), showing that answer-only metrics can overestimate reasoning robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。