SMART为大模型数学推理能力提供多维度评估,发现模型存在隐藏短板。
SMART: Self-Generating and Self-Validating Multi-Dimensional Assessment for LLMs' Mathematical Problem Solving
- 按波利亚理论拆解数学推理为四维度,设计针对性测评任务
- 22个模型测试显示各维度表现差异显著,暴露真实能力短板
- 提出全通过率新指标,更准确衡量综合解题能力
大型语言模型在各类数学基准上表现优异,但其成功是否反映真正推理能力仍存疑。现有评估方法通常只关注最终答案或中间推理步骤,将数学推理简化为单一输入输出映射,忽视其多阶段、多维度的认知本质。受波利亚问题解决理论启发,我们提出SMART基准,将数学问题求解分解为四个认知维度:语义理解、数学推理、算术计算和反思与优化,并设计对应维度的任务来测量模型的相应认知过程。我们对22个最先进的开源与闭源大模型应用SMART,揭示了它们在各维度能力上的显著差异。研究发现当前模型存在真实弱点,并由此提出新指标‘全通过率’(All-Pass Score),以更准确捕捉真正的解题能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks. However, concerns remain as to whether these successes reflect genuine reasoning or superficial pattern recognition. Existing evaluation methods, which typically focus either on the final answer or on the intermediate reasoning steps, reduce mathematical reasoning to a shallow input-output mapping, overlooking its inherently multi-stage and multi-dimensional cognitive nature. Inspired by Polya's problem-solving theory, we propose SMART, a benchmark that decomposes mathematical problem-solving into four cognitive dimensions: Semantic Understanding, Mathematical Reasoning, Arithmetic Computation, and Reflection & Refinement, and introduces dimension-specific tasks to measure the corresponding cognitive processes of LLMs. We apply SMART to 22 state-of-the-art open- and closed-source LLMs and uncover substantial discrepancies in their capabilities across dimensions. Our findings reveal genuine weaknesses in current models and motivate a new metric, the All-Pass Score, designed to better capture true problem-solving capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。