arXiv:2409.13120cs.CLcs.AI2024-09被引 26

测试大模型能否像人一样评分作文,发现它们打分偏高且不靠谱。

Are Large Language Models Good Essay Graders?

  • 用零样本和少样本提示评估ChatGPT与Llama的作文评分能力
  • 大模型评分普遍低于人类,且相关性差,尤其ChatGPT更严厉
  • 常见作文特征如长度、连接词、可读性等与评分无关,适合辅助而非替代

我们评估大型语言模型(LLMs)在作文评分任务中的有效性,重点关注其与人工评分的一致性。具体考察了ChatGPT与Llama在自动作文评分(AES)任务中的表现,这是教育领域重要的自然语言处理应用。采用零样本和少样本学习以及不同提示方法,在ASAP数据集上对比模型给出的分数与人工评分。结果表明,两类模型的评分普遍低于人类评分,且与人类评分相关性较弱;其中ChatGPT比Llama更严苛,偏差更大。我们还检验了以往AES方法中常用的作文特征,包括长度、连接词使用、可读性指标及拼写语法错误数,发现这些特征均未与人工或模型评分形成强相关。最后,报告了Llama 3的结果,整体表现优于前代模型。总体而言,尽管当前大模型尚不足以替代人工评分,但其表现对辅助人类批改作文具有一定的前景。

原文摘要 · Abstract (English)

We evaluate the effectiveness of Large Language Models (LLMs) in assessing essay quality, focusing on their alignment with human grading. More precisely, we evaluate ChatGPT and Llama in the Automated Essay Scoring (AES) task, a crucial natural language processing (NLP) application in Education. We consider both zero-shot and few-shot learning and different prompting approaches. We compare the numeric grade provided by the LLMs to human rater-provided scores utilizing the ASAP dataset, a well-known benchmark for the AES task. Our research reveals that both LLMs generally assign lower scores compared to those provided by the human raters; moreover, those scores do not correlate well with those provided by the humans. In particular, ChatGPT tends to be harsher and further misaligned with human evaluations than Llama. We also experiment with a number of essay features commonly used by previous AES methods, related to length, usage of connectives and transition words, and readability metrics, including the number of spelling and grammar mistakes. We find that, generally, none of these features correlates strongly with human or LLM scores. Finally, we report results on Llama 3, which are generally better across the board, as expected. Overall, while LLMs do not seem an adequate replacement for human grading, our results are somewhat encouraging for their use as a tool to assist humans in the grading of written essays in the future.

作文评分大模型评估教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。