arXiv:2412.18719cs.CL2024-12中稿 · IJAIE被引 32

用大模型自动评阅科学写作,效果媲美教师且更可靠。

Using Large Language Models for Automated Grading of Student Writing about Science

  • 用GPT-4结合评分标准与参考答案评估学生作业。
  • 在120名成人学习者12个问题上,表现接近教师评分。
  • 适合大规模在线课程或非专业科学教学场景使用。

大规模班级中对学生科学写作的评估面临巨大挑战。因此,大多数课程,尤其是科学类课程,通常依赖多选题等客观题型进行评价。随着人工智能的快速发展,大语言模型(LLMs)为自动评估学生写作提供了新可能。本研究使用GPT-4评估了其在三个通过Coursera提供的大型开放在线课程(MOOC)中的表现:天文学、天体生物学以及天文学史与哲学。受试者为成人学习者,共120人回答了12个问题。GPT-4在获得教师提供的总分、标准答案及评分细则后,既评估了其对教师评分的复现可靠性,也尝试生成自己的评分标准。结果显示,该模型在整体和个体层面均优于同伴评分,且在三门课程中与教师评分基本一致。结果表明,大语言模型可实现自动化、可靠且可扩展的学生科学写作评分。

原文摘要 · Abstract (English)

Assessing writing in large classes for formal or informal learners presents a significant challenge. Consequently, most large classes, particularly in science, rely on objective assessment tools such as multiple-choice quizzes, which have a single correct answer. The rapid development of AI has introduced the possibility of using large language models (LLMs) to evaluate student writing. An experiment was conducted using GPT-4 to determine if machine learning methods based on LLMs can match or exceed the reliability of instructor grading in evaluating short writing assignments on topics in astronomy. The audience consisted of adult learners in three massive open online courses (MOOCs) offered through Coursera. One course was on astronomy, the second was on astrobiology, and the third was on the history and philosophy of astronomy. The results should also be applicable to non-science majors in university settings, where the content and modes of evaluation are similar. The data comprised answers from 120 students to 12 questions across the three courses. GPT-4 was provided with total grades, model answers, and rubrics from an instructor for all three courses. In addition to evaluating how reliably the LLM reproduced instructor grades, the LLM was also tasked with generating its own rubrics. Overall, the LLM was more reliable than peer grading, both in aggregate and by individual student, and approximately matched instructor grades for all three online courses. The implication is that LLMs may soon be used for automated, reliable, and scalable grading of student science writing.

自动评分大模型应用教育AI写作评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。