用大模型自动评数学开放题,比谁更准更贴心。
Automated Feedback in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
- 用微调的Mistral和SBERT、零样本的GPT4评估学生解题反馈
- 教师评分显示GPT4在准确性和反馈质量上最优
- 适合教育科技开发者和个性化学习系统研究者
自动化反馈在提升学习成效方面已有充分验证。近年来大型语言模型(LLMs)在增强自动化反馈系统方面展现出潜力。本研究旨在探索LLMs在数学教育中提供自动化反馈的可行性。通过对比Llama、SBERT-Canberra和GPT4三种模型,评估其对开放式数学问题的学生作答进行评分与生成定性反馈的能力。采用专为数学优化的Mistral模型,基于中学生数学题的学生作答与教师反馈数据集进行微调;类似方法用于训练SBERT模型,而GPT4则采用零样本学习方式。通过两位教师依据统一评分量表对生成反馈的准确性与相关性进行评判,开展定量与定性分析。研究结果表明,GPT4在评分准确性和反馈质量方面表现最佳。本研究为自动化反馈系统的发展提供了实证支持,并指出了未来利用生成式大模型实现个性化学习体验的方向。
原文摘要 · Abstract (English)
The effectiveness of feedback in enhancing learning outcomes is well documented within Educational Data Mining (EDM). Various prior research has explored methodologies to enhance the effectiveness of feedback. Recent developments in Large Language Models (LLMs) have extended their utility in enhancing automated feedback systems. This study aims to explore the potential of LLMs in facilitating automated feedback in math education. We examine the effectiveness of LLMs in evaluating student responses by comparing 3 different models: Llama, SBERT-Canberra, and GPT4 model. The evaluation requires the model to provide both a quantitative score and qualitative feedback on the student's responses to open-ended math problems. We employ Mistral, a version of Llama catered to math, and fine-tune this model for evaluating student responses by leveraging a dataset of student responses and teacher-written feedback for middle-school math problems. A similar approach was taken for training the SBERT model as well, while the GPT4 model used a zero-shot learning approach. We evaluate the model's performance in scoring accuracy and the quality of feedback by utilizing judgments from 2 teachers. The teachers utilized a shared rubric in assessing the accuracy and relevance of the generated feedback. We conduct both quantitative and qualitative analyses of the model performance. By offering a detailed comparison of these methods, this study aims to further the ongoing development of automated feedback systems and outlines potential future directions for leveraging generative LLMs to create more personalized learning experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。