用大模型实现几何题自动批改与个性化反馈,支持多次迭代练习
Personalized Auto-Grading and Feedback System for Constructive Geometry Tasks Using Large Language Models on an Online Math Platform
- 基于GPT-4的提示词机制,对比学生答案与标准解法进行评分
- 79名中学生测试显示评分与教师判断高度一致,多数学生通过反馈修正错误
- 适合需要即时反馈的在线数学教学场景,尤其适用于多步骤几何题
随着个性化学习在数学教育中的兴起,对能实时评估复杂学生作答并提供定制化反馈的智能系统需求日益增长。本研究在韩国在线几何工具Algeomath平台上,构建了一个针对构造几何任务的个性化自动批改与反馈系统,利用大语言模型(LLMs)实现。该系统通过分析学生提交的几何构造过程,评估其程序准确性与概念理解水平。采用基于提示的评分机制,结合GPT-4与少样本学习方法,将学生答案与模型解法进行对比。反馈依据教师编写的典型学生回答示例生成,并动态适配学生的解题历史,每道题支持最多四次迭代尝试。在79名初中生的小规模试点中,系统生成的评分与教师评判高度一致,反馈帮助多数学生修正错误并完成多步几何任务。尽管短期纠错效果明显,长期迁移效果尚不明确。整体表明,LLM具备支持可扩展、教师对齐的形成性评价潜力,但术语处理与反馈设计仍有优化空间。
原文摘要 · Abstract (English)
As personalized learning gains increasing attention in mathematics education, there is a growing demand for intelligent systems that can assess complex student responses and provide individualized feedback in real time. In this study, we present a personalized auto-grading and feedback system for constructive geometry tasks, developed using large language models (LLMs) and deployed on the Algeomath platform, a Korean online tool designed for interactive geometric constructions. The proposed system evaluates student-submitted geometric constructions by analyzing their procedural accuracy and conceptual understanding. It employs a prompt-based grading mechanism using GPT-4, where student answers and model solutions are compared through a few-shot learning approach. Feedback is generated based on teacher-authored examples built from anticipated student responses, and it dynamically adapts to the student's problem-solving history, allowing up to four iterative attempts per question. The system was piloted with 79 middle-school students, where LLM-generated grades and feedback were benchmarked against teacher judgments. Grading closely aligned with teachers, and feedback helped many students revise errors and complete multi-step geometry tasks. While short-term corrections were frequent, longer-term transfer effects were less clear. Overall, the study highlights the potential of LLMs to support scalable, teacher-aligned formative assessment in mathematics, while pointing to improvements needed in terminology handling and feedback design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。