arXiv:2504.12347cs.CLcs.AI2025-04

评测大模型数学能力,发现其水平已接近顶尖学生。

Assessment of Evolving Large Language Models in Upper Secondary Mathematics

  • 用芬兰高考题测试大模型数学能力
  • 部分模型得分接近或达到满分
  • 适合教育科技研究者与教学工具开发者

大型语言模型(LLMs)在教育场景中展现出日益增长的潜力,但其数学推理能力仍处于发展之中。本研究利用芬兰高中毕业考试——一项针对高中教育的高利害数字化考试,评估了多种大模型的数学能力。初步测试显示模型表现中等,对应中等成绩;但后续评估表明,随着模型迭代,其表现显著提升。值得注意的是,部分模型取得了近乎完美或满分的成绩,达到顶尖学生水平,并具备大学入学资格。研究结果凸显了大模型数学能力的快速进步,也展示了其作为学习与教学支持工具的潜在价值。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown increasing promise in educational settings, yet their mathematical reasoning has been considered evolving. This study evaluates the mathematical capabilities of various LLMs using the Finnish matriculation examination, a high-stakes digital test for upper secondary education. Initial tests yielded moderate performance corresponding to mid-range grades, but later evaluations demonstrated substantial improvements as the language models evolved. Remarkably, some models achieved near-perfect or perfect scores, matching top student performance and qualifying for university admission. Our findings highlight the rapid advances in the mathematical proficiency of LLMs and illustrate their potential as underlying tools to support learning and teaching in a variety of ways.

大模型数学推理教育应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。