用大模型自动批改生物信息学作业,效果接近人工且更私密。
Automated Assignment Grading with Large Language Models: Insights From a Bioinformatics Course
- 用精心设计的提示词让大模型自动批改文本作业。
- 盲测显示大模型反馈质量与人工相当,准确率超90%。
- 开源模型表现媲美商业模型,适合注重隐私的学校使用。
为学生提供个性化反馈是教育的核心,能显著提升学习效果。然而,在大规模班级中实现个性化反馈因耗时耗力而难以实施。近年来,自然语言处理和大语言模型(LLMs)的发展为这一问题提供了新方案,可在降低教师负担的同时提升学生满意度和学习成效。我们对卢布尔雅那大学2024/25年《生物信息学导论》课程中基于LLM的自动批改系统进行了实际评估。本学期超过100名学生完成了36道文本类题目,大部分由大模型自动评分。在盲测中,学生收到由大模型和人类助教提供的反馈但不知来源,并对反馈质量进行评分。我们系统评估了六种商用及开源大模型,并与人工评分结果对比。结果显示,通过优化提示词,大模型可达到与人类评阅者相当的评分准确率与反馈质量。同时,开源模型表现不逊于商业模型,使学校能在保障数据隐私的前提下自建评分系统。
原文摘要 · Abstract (English)
Providing students with individualized feedback through assignments is a cornerstone of education that supports their learning and development. Studies have shown that timely, high-quality feedback plays a critical role in improving learning outcomes. However, providing personalized feedback on a large scale in classes with large numbers of students is often impractical due to the significant time and effort required. Recent advances in natural language processing and large language models (LLMs) offer a promising solution by enabling the efficient delivery of personalized feedback. These technologies can reduce the workload of course staff while improving student satisfaction and learning outcomes. Their successful implementation, however, requires thorough evaluation and validation in real classrooms. We present the results of a practical evaluation of LLM-based graders for written assignments in the 2024/25 iteration of the Introduction to Bioinformatics course at the University of Ljubljana. Over the course of the semester, more than 100 students answered 36 text-based questions, most of which were automatically graded using LLMs. In a blind study, students received feedback from both LLMs and human teaching assistants without knowing the source, and later rated the quality of the feedback. We conducted a systematic evaluation of six commercial and open-source LLMs and compared their grading performance with human teaching assistants. Our results show that with well-designed prompts, LLMs can achieve grading accuracy and feedback quality comparable to human graders. Our results also suggest that open-source LLMs perform as well as commercial LLMs, allowing schools to implement their own grading systems while maintaining privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。