arXiv:2511.10819cs.CL2025-11被引 2

用GPT-4o给本科语言学作业打分,与人工评分高度一致。

LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation

  • 用GPT-4o自动批改短答和报告,对比人工评分。
  • 短答评分相关性达0.98,55%完全一致;报告整体匹配度高。
  • 适合想探索自动化评分的教育研究者,尤其关注实操效果。

大型语言模型(LLMs)在教育任务如评分中的应用日益增多,但其在真实课堂中与人类评价的一致性仍缺乏深入研究。本研究探讨使用GPT-4o评估本科生计算语言学课程的短答测验和项目报告的可行性。我们收集了约50名学生在五次测验中的作答,以及14支团队的项目报告。将GPT-4o生成的分数与课程助教独立完成的人工评分进行比较。结果表明,GPT-4o与人类评分者相关性高达0.98,在55%的测验案例中实现完全一致得分。对于项目报告,整体评分与人工评分高度一致,但在技术性、开放性回答上存在一定程度的评分波动。本文公开全部代码和样例数据,以支持后续在教育评估中对LLM的研究。该工作揭示了基于LLM评分系统的潜力与局限,推动其在真实学术场景中的发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly explored for educational tasks such as grading, yet their alignment with human evaluation in real classrooms remains underexamined. In this study, we investigate the feasibility of using an LLM (GPT-4o) to evaluate short-answer quizzes and project reports in an undergraduate Computational Linguistics course. We collect responses from approximately 50 students across five quizzes and receive project reports from 14 teams. LLM-generated scores are compared against human evaluations conducted independently by the course teaching assistants (TAs). Our results show that GPT-4o achieves strong correlation with human graders (up to 0.98) and exact score agreement in 55\% of quiz cases. For project reports, it also shows strong overall alignment with human grading, while exhibiting some variability in scoring technical, open-ended responses. We release all code and sample data to support further research on LLMs in educational assessment. This work highlights both the potential and limitations of LLM-based grading systems and contributes to advancing automated grading in real-world academic settings.

自动评分GPT-4o教育AILLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。