arXiv:2509.01395cs.CLcs.AI2025-09EMNLP被引 6

大模型看不懂数学解题错误,即使能看正确答案。

LLMs cannot spot math errors, even when allowed to peek into the solution

  • 通过生成中间修正解,让模型更贴近学生原思路。
  • 在两个数据集上,顶尖模型定位首错步骤准确率仍很低。
  • 适合研究数学推理错误检测与模型可解释性的读者。

大型语言模型(LLMs)在数学应用题上表现优异,但在识别学生解答中的错误等元推理任务中却表现不佳。本文针对逐步求解过程中的首错步骤定位问题,使用两个错误推理数据集(VtG 和 PRM800K)进行实验。结果表明,即使允许模型查看参考答案,当前最先进的大模型依然难以准确找出第一个错误步骤。为此,我们提出一种新方法:生成一个与原学生解法更接近的中间修正解,从而提升定位性能。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate remarkable performance on math word problems, yet they have been shown to struggle with meta-reasoning tasks such as identifying errors in student solutions. In this work, we investigate the challenge of locating the first error step in stepwise solutions using two error reasoning datasets: VtG and PRM800K. Our experiments show that state-of-the-art LLMs struggle to locate the first error step in student solutions even when given access to the reference solution. To that end, we propose an approach that generates an intermediate corrected student solution, aligning more closely with the original student's solution, which helps improve performance.

数学推理错误检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。