arXiv:2606.24973cs.CLcs.AI2026-06

测试大模型对英国中学生真实答卷的评分一致性,效果优于人工评卷。

LLM Performance on a Real, Double-Marked GCSE Benchmark

论文配图:LLM Performance on a Real, Double-Marked GCSE Benchmark
图 1 · 摘自论文原文
  • 构建32,534份双评真实学生答卷数据集,覆盖五科328题。
  • 顶尖模型评分与考官共识一致率高于考官间一致性。
  • 适用于主观题与手写数学题的低成本自动评分场景。

我们引入了一个包含32,534份双评真实学生答卷的数据集,涵盖英国16岁学生参加的GCSE模拟考试,涉及五个学科共328道题目,包括手写内容。测试现成大语言模型与考官的一致性是否达到考官间的相互一致性水平。结果表明,模型在各科目上普遍与考官共识高度一致,表现最佳的模型甚至比考官之间的认同度更高。模型在英语作文等主观任务上得分优异,也能有效处理复杂且混乱的手写数学试卷。评分一致性在考官评分线附近保持稳定,且不随模型规模显著变化,证明其具备成本效益的自动化评分潜力。

原文摘要 · Abstract (English)

We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning 328 questions across five subjects and including handwritten work. We test whether off-the-shelf large language models agree with examiners as closely as the two examiners agree with each other. We find that models overwhelmingly agree well with the examiner consensus across subjects, with the top performing models agreeing more closely with examiners than examiners agree with each other. Models achieve high scores for subjective tasks like English essay marking, as well as handling complex and messy handwritten Maths paper scripts. Agreement is uniform near the examiner line, and not massively discriminated by model size, providing cost-effective automated marking solutions.

自动评分大模型评估教育应用手写识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。