arXiv:2506.24006cs.CLmath.HO2025-06综述被引 5

大模型解题快但不懂现实背景,教学应用有局限。

Large Language Models Don't Make Sense of Word Problems. A Scoping Review from a Mathematics Education Perspective

  • 对比学生与大模型解题思路,发现模型只做表面计算
  • 测试显示模型在287道题中近全对,但真实情境题仍出错
  • 适合用于机械训练,不适合培养数学理解力

大型语言模型(如ChatGPT)在教育中的应用前景引发关注,尤其在数学问题求解方面。尽管它们能处理文本输入,但能否真正理解问题的现实语境仍不明确。本文从数学教育视角开展系统性综述,包含三部分:技术概览、文献梳理及实证评估。技术层面指出,计算机科学中“数学推理”概念与数学教育中的理解存在差异。文献分析213项研究发现,主流数据集以s-问题为主,无需考虑真实情境。实证评估对GPT-3.5-turbo、GPT-4o-mini、GPT-4.1、o3和GPT-5在287道题目上的表现进行测试,结果显示多数模型在简单问题上接近满分,包括在PISA的20道题中得满分;但在涉及非现实或矛盾情境的问题上仍表现不佳。综合三方面结论表明,大模型掌握的是表层解题流程,而非真正理解问题,这限制了其作为教学工具在课堂中的实际价值。

原文摘要 · Abstract (English)

The progress of Large Language Models (LLMs) like ChatGPT raises the question of how they can be integrated into education. One hope is that they can support mathematics learning, including word-problem solving. Since LLMs can handle textual input with ease, they appear well-suited for solving mathematical word problems. Yet their real competence, whether they can make sense of the real-world context, and the implications for classrooms remain unclear. We conducted a scoping review from a mathematics-education perspective, including three parts: a technical overview, a systematic review of word problems used in research, and a state-of-the-art empirical evaluation of LLMs on mathematical word problems. First, in the technical overview, we contrast the conceptualization of word problems and their solution processes between LLMs and students. In computer-science research this is typically labeled mathematical reasoning, a term that does not align with usage in mathematics education. Second, our literature review of 213 studies shows that the most popular word-problem corpora are dominated by s-problems, which do not require a consideration of realities of their real-world context. Finally, our evaluation of GPT-3.5-turbo, GPT-4o-mini, GPT-4.1, o3, and GPT-5 on 287 word problems shows that most recent LLMs solve these s-problems with near-perfect accuracy, including a perfect score on 20 problems from PISA. LLMs still showed weaknesses in tackling problems where the real-world context is problematic or non-sensical. In sum, we argue based on all three aspects that LLMs have mastered a superficial solution process but do not make sense of word problems, which potentially limits their value as instructional tools in mathematics classrooms.

大模型数学教育解题能力s-问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。