arXiv:2503.16460cs.HCcs.AI2025-03被引 36

评测大模型在数学辅导中的表现,发现其解题正确但教学易出错。

Beyond Final Answers: Evaluating Large Language Models for Math Tutoring

  • 用大学代数系统生成测试题,对比大模型解题与标准答案。
  • 大模型解题正确率达85.5%,但作为导师时仅56.6%全程无误。
  • 适合需要人工校验的数学辅导场景,不建议直接替代真人教师。

尽管大型语言模型(LLMs)在解决数学问题方面取得显著进展,如GSM8k、ProofNet、AlphaGeometry和MathOdyssey所示,其在数学辅导场景中的可靠性仍待深入研究。本文提出两种新方法评估LLMs在数学辅导中的正确性与教学质量:第一种以大学代数智能辅导系统为测试平台,生成基准题目,让多种LLMs求解,并与系统生成解答对比;第二种则将人类评估者作为学生,体验与不同LLMs的互动辅导过程,通过定性编码评估其教学质量和正确性。实验涵盖ChatGPT系列模型(3.5 Turbo、4、4o、o1-mini、o1-preview)。结果显示,作为解题工具时,大模型对85.5%的题目给出正确最终答案;作为交互式导师时,90%的对话提供高质量教学支持,但仅56.6%完全正确。结论指出,当前大模型尚不足以独立承担数学智能辅导任务,需人工监督或额外纠错机制保障质量。

原文摘要 · Abstract (English)

Researchers have made notable progress in applying Large Language Models (LLMs) to solve math problems, as demonstrated through efforts like GSM8k, ProofNet, AlphaGeometry, and MathOdyssey. This progress has sparked interest in their potential use for tutoring students in mathematics. However, the reliability of LLMs in tutoring contexts -- where correctness and instructional quality are crucial -- remains underexplored. Moreover, LLM problem-solving capabilities may not necessarily translate into effective tutoring support for students. In this work, we present two novel approaches to evaluate the correctness and quality of LLMs in math tutoring contexts. The first approach uses an intelligent tutoring system for college algebra as a testbed to assess LLM problem-solving capabilities. We generate benchmark problems using the tutor, prompt a diverse set of LLMs to solve them, and compare the solutions to those generated by the tutor. The second approach evaluates LLM as tutors rather than problem solvers. We employ human evaluators, who act as students seeking tutoring support from each LLM. We then assess the quality and correctness of the support provided by the LLMs via a qualitative coding process. We applied these methods to evaluate several ChatGPT models, including 3.5 Turbo, 4, 4o, o1-mini, and o1-preview. Our findings show that when used as problem solvers, LLMs generate correct final answers for 85.5% of the college algebra problems tested. When employed interactively as tutors, 90% of LLM dialogues show high-quality instructional support; however, many contain errors -- only 56.6% are entirely correct. We conclude that, despite their potential, LLMs are not yet suitable as intelligent tutors for math without human oversight or additional mechanisms to ensure correctness and quality.

数学辅导大模型评测教育AI对话质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。