arXiv:2603.25633cs.AI2026-03

大模型解题越准,评估错误步骤也越准,但诊断仍需额外能力。

Is Mathematical Problem-Solving Expertise in Large Language Models Associated with Assessment Performance?

  • 用GPT-4和GPT-5在GSM8K与MATH上测试解题与评估能力
  • 解对的题目,评估准确率显著高于解错的题目
  • 适合关注AI数学教学评估系统设计的研究者

大型语言模型(LLMs)在数学教育中不仅用于解题,还被用于评估学习者的推理过程。然而,解题能力强是否意味着评估能力也强尚不明确。本研究基于PROCESSBENCH中的GSM8K和MATH子集,使用人类标注的基准数据,考察两个基于LLM的数学辅导代理(分别采用GPT-4和GPT-5)在相同数学问题上的两种任务表现:求解原始问题和预测基准提供的解答中最早出错的步骤。结果显示,在同一模型内部存在一致模式:对于该模型正确求解的问题,其评估准确率显著高于模型求解错误的问题,且在两个模型和数据集上均呈统计学显著相关。同时,评估任务整体难度高于直接求解,尤其在包含错误的解答中更为明显。结果表明,数学解题专长有助于提升评估表现,但可靠的逐步诊断还需额外能力,如步骤追踪、监控与精确误差定位。这些发现对人工智能支持的自适应教学系统(AISs)在数学形成性评估中的设计与评估具有启示意义。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used in math education not only as problem solvers but also as assessors of learners' reasoning. However, it remains unclear whether stronger math problem-solving ability is associated with stronger step-level assessment performance. This study examines that relationship using the GSM8K and MATH subsets of PROCESSBENCH, a human-annotated benchmark for identifying the earliest erroneous step in mathematical reasoning. We evaluate two LLM-based math tutor agent settings, instantiated with GPT-4 and GPT-5, in two independent tasks on the same math problems: solving the original problem and assessing a benchmark-provided solution by predicting the earliest erroneous step. Results show a consistent within-model pattern: assessment accuracy is substantially higher on math problem items the same model solved correctly than on items it solved incorrectly, with statistically significant associations across both models and datasets. At the same time, assessment remains more difficult than direct problem solving, especially on error-present solutions. These findings suggest that math problem-solving expertise supports stronger assessment performance, but reliable step-level diagnosis also requires additional capabilities such as step tracking, monitoring, and precise error localization. The results have implications for the design and evaluation of AI-supported Adaptive Instructional Systems (AISs) for formative assessment in math education.

数学推理AI评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。