arXiv:2601.16314cs.CLcs.AI2026-01被引 3

用大模型自动评全国毕业作文,效果接近人工且可生成个性化反馈

Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP

  • 基于课程标准构建评分规则,结合大模型与统计NLP进行自动评分
  • 自动评分结果与人工评分高度一致,误差在人类打分范围内
  • 适合需要大规模电子化考试的国家,尤其小语种地区可借鉴

大型语言模型(LLMs)可实现对开放性答题的快速、一致自动化评分,涵盖内容与论证等传统需人工判断维度。本文在爱沙尼亚两个完整国家级考前作文数据集上检验了自动化评分的适用性。通过将官方课程评分标准操作化,对比了基于LLM和统计自然语言处理(NLP)的评分与人工评分小组结果。结果显示,自动化评分表现可媲美人类评审,且多数评分落在人类评分区间内。我们还评估了偏见、提示注入风险以及大模型作为写作者的能力。研究证明,以评分标准为驱动、人类参与的评分流程可在高风险写作评估中落地,特别适用于即将全面推行电子化考试的数字化社会如爱沙尼亚。此外,系统可生成细粒度子分报告,用于提供系统性、个性化的教学与备考反馈。该研究为在小语种背景下实现国家级自动化评估提供了实证支持,同时保障了人类监督与新兴教育监管标准合规。

原文摘要 · Abstract (English)

Large language models (LLMs) enable rapid and consistent automated evaluation of open-ended exam responses, including dimensions of content and argumentation that have traditionally required human judgment. This is particularly important in cases where a large amount of exams need to be graded in a limited time frame, such as nation-wide graduation exams in various countries. Here, we examine the applicability of automated scoring on two large datasets of trial exam essays of two full national cohorts from Estonia. We operationalize the official curriculum-based rubric and compare LLM and statistical natural language processing (NLP) based assessments with human panel scores. The results show that automated scoring can achieve performance comparable to that of human raters and tends to fall within the human scoring range. We also evaluate bias, prompt injection risks, and LLMs as essay writers. These findings demonstrate that a principled, rubric-driven, human-in-the-loop scoring pipeline is viable for high-stakes writing assessment, particularly relevant for digitally advanced societies like Estonia, which is about to adapt a fully electronic examination system. Furthermore, the system produces fine-grained subscore profiles that can be used to generate systematic, personalized feedback for instruction and exam preparation. The study provides evidence that LLM-assisted assessment can be implemented at a national scale, even in a small-language context, while maintaining human oversight and compliance with emerging educational and regulatory standards.

自动评分大模型教育评估小语种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。