构建数学能力评估基准,提升大模型数值推理表现
PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

- 设计分层多模态数学题库,覆盖4类认知能力
- 引入SOLVE与IRPO模块,数值计算准确率显著提升
- 适合关注数学推理与工具调用的大模型研究者
尽管数值推理是大语言模型在各类应用中数学能力的核心,但现有评估基准普遍缺乏对数值处理与数学推理的融合考量,难以解释模型在数学任务中的失败原因。我们提出PyraMathBench,一个涵盖32,505道题目、来自7,404个数学应用题的分层综合性评估基准,覆盖4个关键认知维度、14个子类别及2种模态。实验表明,大模型在数值计算不足和抽象数字问题处理方面表现严重受限。为此,我们提出Smart Optimization & Learning-based VErsatile模块(SOLVE)与交互式相对策略优化(IRPO),通过高效工具调用(模糊匹配与低质量调用拒绝)增强模型的数值-数学协同能力。对比实验显示,Qwen-2.5在采用SOLVE与IRPO训练后,得分提升5.0分。
原文摘要 · Abstract (English)
Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and mathematical reasoning, hindering the interpretability of failures in math tasks. We introduce PyraMathBench, a comprehensive hierarchical benchmark with 32,505 questions derived from 7,404 math word problems, spanning 4 key cognitive aspects, 14 subcategories, and 2 modalities. Experiments reveal that LLMs' performance is severely compromised by inadequate numerical computation and weak handling of abstract numerical questions. To address this, we propose the Smart Optimization & Learning-based VErsatile module (SOLVE) and Interactive Relative Policy Optimization (IRPO), which enhance LLMs' numerical-mathematical synergy via efficient tool calls (fuzzy matching and low-quality call rejection). Comparative experiments show Qwen-2.5 achieves a 5.0 score improvement with SOLVE and IRPO training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。