arXiv:2411.05231cs.CYcs.CL2024-11被引 14

用GPT-4o自动批改大学数学手写答题,效果仍有提升空间。

Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams

  • 结合视觉与文本信息,用GPT-4o分析手写数学答案。
  • 在概率论考试中,评分与人工打分的对齐度不足。
  • 提示词设计影响评分准确率,需进一步优化。

生成式人工智能在评估开放式学生作答方面展现出潜力,但因数据匮乏及视觉与文本信息融合困难,极少研究涉及手写答案的评分。本文利用先进的多模态AI模型(特别是GPT-4o),自动批改大学水平数学考试的手写解答。基于真实学生在概率论考试中的答题数据,我们通过多种提示策略评估GPT-4o与人工评分标准的一致性。结果表明,虽提供评分细则能提升对齐度,但整体准确率仍不足以满足实际应用需求,说明该任务仍有显著提升空间。

原文摘要 · Abstract (English)

Recent advances in generative artificial intelligence (AI) have shown promise in accurately grading open-ended student responses. However, few prior works have explored grading handwritten responses due to a lack of data and the challenge of combining visual and textual information. In this work, we leverage state-of-the-art multi-modal AI models, in particular GPT-4o, to automatically grade handwritten responses to college-level math exams. Using real student responses to questions in a probability theory exam, we evaluate GPT-4o's alignment with ground-truth scores from human graders using various prompting techniques. We find that while providing rubrics improves alignment, the model's overall accuracy is still too low for real-world settings, showing there is significant room for growth in this task.

AI批改多模态数学教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。