用大模型自动批改西语开放题,准确率超95%。
On the effectiveness of LLMs for automatic grading of open-ended questions in Spanish
- 测试多种大模型与提示词风格,聚焦西语答题场景。
- 三等级评分准确率超95%,二分类任务达98%以上。
- 提示词设计影响显著,适合教育自动化落地应用。
批改是教师面临的一项耗时且繁重的任务,对学习者反馈至关重要,及时反馈可提升学习效果。近年来,大语言模型(LLMs)在自动批改方面展现出潜力。本文研究不同大模型及提示技术在自动批改西语简答开放题中的表现。与多数文献不同,本研究关注问题、答案和提示均使用西班牙语的实际应用场景。实验结果表明,先进大模型(包括开源与专有模型)在与人工专家评分对比中,在准确性、精确性和一致性方面表现良好。结果对提示风格敏感,提示中特定词汇或内容可能引入偏差。但最优模型与提示组合在三等级评分任务中准确率持续超过95%,在简化为对错二分类任务时更达98%以上,充分展现了大模型在教育自动化中的应用潜力。
原文摘要 · Abstract (English)
Grading is a time-consuming and laborious task that educators must face. It is an important task since it provides feedback signals to learners, and it has been demonstrated that timely feedback improves the learning process. In recent years, the irruption of LLMs has shed light on the effectiveness of automatic grading. In this paper, we explore the performance of different LLMs and prompting techniques in automatically grading short-text answers to open-ended questions. Unlike most of the literature, our study focuses on a use case where the questions, answers, and prompts are all in Spanish. Experimental results comparing automatic scores to those of human-expert evaluators show good outcomes in terms of accuracy, precision and consistency for advanced LLMs, both open and proprietary. Results are notably sensitive to prompt styles, suggesting biases toward certain words or content in the prompt. However, the best combinations of models and prompt strategies, consistently surpasses an accuracy of 95% in a three-level grading task, which even rises up to more than 98% when the it is simplified to a binary right or wrong rating problem, which demonstrates the potential that LLMs have to implement this type of automation in education applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。