arXiv:2603.23714cs.AIcs.CL2026-03被引 1

LLM评分与人类差异大,短文偏高分,长文错字就扣分。

LLMs Do Not Grade Essays Like Humans

  • 直接用GPT和Llama模型评分,不微调
  • 短文得分偏高,长文有小错就被压分
  • 评分与反馈一致,适合辅助打分

大型语言模型被提议用于自动作文评分,但其与人工评分的一致性尚不明确。本文评估了GPT和Llama系列模型在无需任务特训的零样本设置下生成的评分与人工评分的对比情况。结果表明,模型评分与人工评分的吻合度较弱,且受作文特征影响。相比人类评分员,模型更倾向于给短篇或内容不充分的作文打高分,而对较长但存在轻微语法或拼写错误的作文打低分。此外,模型生成的评分与其反馈高度一致:获得较多表扬的作文得分更高,批评较多的得分更低。这说明模型评分与反馈具有内在一致性,但依赖的判断信号与人类不同,导致与人工评分实践存在显著偏差。尽管如此,本研究显示模型生成的反馈与评分可保持稳定一致,具备作为作文评分辅助工具的潜力。

原文摘要 · Abstract (English)

Large language models have recently been proposed as tools for automated essay scoring, but their agreement with human grading remains unclear. In this work, we evaluate how LLM-generated scores compare with human grades and analyze the grading behavior of several models from the GPT and Llama families in an out-of-the-box setting, without task-specific training. Our results show that agreement between LLM and human scores remains relatively weak and varies with essay characteristics. In particular, compared to human raters, LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors. We also find that the scores generated by LLMs are generally consistent with the feedback they generate: essays receiving more praise tend to receive higher scores, while essays receiving more criticism tend to receive lower scores. These results suggest that LLM-generated scores and feedback follow coherent patterns but rely on signals that differ from those used by human raters, resulting in limited alignment with human grading practices. Nevertheless, our work shows that LLMs produce feedback that is consistent with their grading and that they can be reliably used in supporting essay scoring.

自动评分LLM评估教育技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。