arXiv:2606.10254cs.AIcs.CL2026-06

评测大模型对真实学生数学推理的判断能力,发现其表现远不如预期。

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

论文配图:RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
图 1 · 摘自论文原文
  • 构建224份真实学生答题数据集,用于评估大模型的评判能力。
  • 模型在真实学生答案上误差达2.96,远高于对合成答案的1.17。
  • 真实学生错误模式复杂多样,现有模型难以适应,适合教育评估研究者。

尽管大型语言模型在解答高中数学题上已接近完美,但其对真实学生多样化推理过程的评估能力仍缺乏研究。为此,我们提出「RealMath-Eval」,一个经过严格标注的224份真实高中考试作答组成的基准数据集。初步评估显示,即使最先进的LLM裁判在该任务上表现显著不足,与专家评分相比均方误差高达~2.96。我们进一步对比了模型在合成的LLM生成解法上的表现,发现其误差仅为~1.17。分析表明,模型在真实学生答案上存在明显“评估差距”:合成错误趋于结构化、低维线性分布,而真实错误空间更丰富;信息论探测显示学生推理具有更高不可预测性,属于当前模型的长尾分布。此外,仅靠表面风格迁移无法弥合差距。结果表明,依赖合成数据的现有评估流程可能无法充分捕捉真实学生的数学推理多样性。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined. To bridge this gap, we introduce \textbf{RealMath-Eval}, a rigorously annotated benchmark of 224 real-world exam responses from high schools. Our initial evaluation reveals that even state-of-the-art LLM judges struggle significantly on this task, exhibiting a high Mean Squared Error ($\sim$2.96) against expert human grading. To probe a plausible explanation, we contrast this performance with a control setting where the same judges evaluate synthetic LLM-generated solutions. We identify a stark ``Evaluation Gap'': judges are considerably more accurate and consistent on synthetic text (MSE $\sim$1.17) but struggle to generalize to authentic student reasoning. Through semantic embedding analysis, we find that synthetic errors suffer from a ``structural collapse'' into predictable, low-dimensional linear subspaces, whereas human errors form a more diverse error space. Furthermore, generative probability probes suggest that human reasoning involves significantly higher information-theoretic surprisal, indicating that student reasoning transitions are more out-of-distribution for current models. Finally, we find that surface-level style transfer fails to close this gap. Our findings suggest that current LLM evaluation pipelines relying heavily on synthetic data may not adequately capture the diversity of authentic student mathematical reasoning.

大模型评估数学推理教育AI真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。