综合65项研究发现,大模型与人工评分一致性受情境影响极大。
Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis
- 按PRISMA 2020指南整合65项研究,覆盖2022年1月至2025年8月
- 大模型与人工评分一致性在不同研究间差异显著,无统一标准
- 结果表明评分一致性高度依赖具体语境,非普遍适用
尽管大型语言模型(LLMs)在自动作文评分(AES)中展现出巨大潜力,但其与人工评分者之间一致性的实证研究结果仍不一致。本研究遵循PRISMA 2020指南,整合了2022年1月至2025年8月期间发表和未发表的65项研究,分析了LLM生成评分与人工评分之间的吻合度。结果显示,一致性水平在不同研究间及同一研究内均存在显著差异,数值跨度广泛。总体而言,研究表明大模型与人工评分的一致性具有高度情境依赖性。论文进一步讨论了其影响、挑战及未来研究方向。
原文摘要 · Abstract (English)
Despite the growing promise of large language models (LLMs) in automated essay scoring (AES), empirical findings regarding their reliability compared to human raters remain mixed. Following the PRISMA 2020 guidelines, we synthesized 65 published and unpublished studies from January 2022 to August 2025 that examined agreement between LLM-generated scores and human ratings. Agreement levels varied substantially both across and within studies, with reported values spanning a wide range. Overall, the findings suggest that LLM-human agreement is highly context-dependent. Implications, challenges, and directions for future research are discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。