AI可辅助批改作文,尤其在语言表达上表现接近教师水平。
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
- 用5个大模型对比教师评分,评估德国中学生作文
- GPT-o1模型与教师评分相关性达0.74,内部一致性达0.80
- 适合想减轻批改负担的教育工作者参考
人工批改学生作文耗时费力,而生成式AI如大语言模型(LLMs)有望缓解这一压力。本研究评估了开源与闭源大模型在评估德国七、八年级学生作文中的表现,对比37位教师对10项指标(如情节逻辑、表达)的评分。使用包含20篇真实作文的语料库,测试了GPT-3.5、GPT-4、o1、LLaMA 3-70B和Mixtral 8x7B五款模型。结果显示,闭源模型整体优于开源模型,尤其在语言类指标上表现突出;其中新型o1模型在总分上与人类评分的斯皮尔曼相关系数达到0.74,内部一致性(ICC)为0.80。表明大模型可有效辅助作文评分,尤其在语言层面,但因倾向于给出更高分数,仍需优化以更好捕捉内容质量。
原文摘要 · Abstract (English)
The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI, such as large language models, offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance and reliability of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs' scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The novel o1 model outperforms all other LLMs, achieving Spearman's $r = .74$ with human assessments in the overall score, and an internal consistency of $ICC=.80$. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency for higher scores, the models require further refinement to better capture aspects of content quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。