构建数学解题过程反馈数据集,评估AI生成反馈的可靠性与教学适配性。
MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support
- 构建包含学生解题过程与教师反馈的MathEDU数据集
- 微调可提升正确性判断与错误定位准确率,但反馈质量仍不足
- 生成反馈常冗长且无法精准揭示学生认知误区,需更懂教育的AI
大型语言模型(LLMs)在教育领域的应用日益广泛,学生逐渐将其作为学习工具。尽管已有研究考察了LLMs的数学能力,但其在真实学生解题过程评分及有效反馈生成方面的可靠性仍待深入探索。本研究提出MathEDU数据集,包含学生数学解题过程与对应教师反馈。我们系统评估多种模型在三个层级任务中的表现:答案正确性分类、错误识别与反馈生成。实验表明,微调策略能有效提升正确性分类与错误步骤定位性能;然而,各模型生成的反馈与教师反馈存在显著差距,主要表现为内容冗长、缺乏针对性,未能精准揭示学生深层误解。这凸显了构建可信且具备教学意识的AI反馈系统的紧迫性。
原文摘要 · Abstract (English)
The increasing reliance on Large Language Models (LLMs) across various domains extends to education, where students progressively use generative AI as a tool for learning. While prior work has examined LLMs' mathematical ability, their reliability in grading authentic student problem-solving processes and delivering effective feedback remains underexplored. This study introduces MathEDU, a dataset consisting of student problem-solving processes in mathematics and corresponding teacher-written feedback. We systematically evaluate the reliability of various models across three hierarchical tasks: answer correctness classification, error identification, and feedback generation. Experimental results show that fine-tuning strategies effectively improve performance in classifying correctness and locating erroneous steps. However, the generated feedback across models shows a considerable gap from teacher-written feedback. Critically, the generated feedback is often verbose and fails to provide targeted explanations for the student's underlying misconceptions. This emphasizes the urgent need for trustworthy and pedagogy-aware AI feedback in education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。