LLM评分有效性取决于题目类型,而非模型本身。
LLM-as-a-judge validity in physics assessment depends more on the task than the model
- 按题目结构区分评分效果,结构化题更可靠。
- 论文题和图表题中模型排名一致率低,评分偏差大。
- 适合结构化试题评估,不适用于开放性写作评分。
随着大语言模型(LLMs)在自动化评估与反馈中的应用增多,理解其作为评分者(LLM-as-a-judge)的有效性至关重要。本文在三类物理测评形式——结构化问题、书面论述和科学图表——上,对比GPT-5.2、Grok 4.1、Claude Opus 4.5、DeepSeek-V3.2、Gemini Pro 3及委员会集成结果与人工评分者的差异,涵盖盲评、提供解法、虚假解法和锚定范例等条件。研究区分绝对准确性和排序一致性:系统可匹配人类评分分布,但无法正确排序作答质量。结果显示,对于771份盲评大学试题和1151份中学及大学结构化题,模型与人类的排序一致性(斯皮尔曼相关系数ρ > 0.6)良好,且官方解法可降低误差并增强一致性;虚假解法虽损害绝对准确性,但排序未受影响。而在55份论文脚本(共275篇)中,盲评下模型评分更严且波动更大,提供评分标准也无法提升排序一致性。锚定范例使模型平均分接近人类,方差低于人类,但排序一致性仍趋近于零。对于1400个代码生成的图表元素,模型排序一致性高(ρ > 0.84),校准接近线性。总体而言,评估有效性主要由任务结构决定,即分数能否映射到明确可观测的评分特征,以及人类基准的可靠性,而非模型本身能力。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge marking across three physics assessment formats - structured questions, written essays, and scientific plots - comparing GPT-5.2, Grok 4.1, Claude Opus 4.5, DeepSeek-V3.2, Gemini Pro 3, and committee aggregations against human markers under blind, solution-provided, false-solution, and anchored conditions. We distinguish absolute accuracy from rank-order agreement, since a marking system can match the distribution of human marks while failing to order responses by quality. Across task types, performance is sharply task-dependent. For blind university exam questions ($n=771$) and secondary and university structured questions ($n=1151$), models show robust rank-order agreement with human markers (Spearman $ρ> 0.6$), with official solutions reducing error and strengthening agreement. False solutions degrade absolute accuracy, showing that models defer to provided references, but leave rank-ordering intact. Essay marking behaves fundamentally differently. Across $n=55$ scripts ($n=275$ essays), blind AI marking is harsher and more variable than human marking and adding a mark scheme does not improve rank-order agreement. Anchored exemplars shift the AI mean close to the human mean and compress variance below the human standard deviation, but rank-order agreement remains near-zero. For code-based plot elements ($n=1400$), models achieve high rank-order agreement ($ρ> 0.84$) with near-linear calibration. Across all task types, validity tracks the structure of the assessment task - the extent to which marks can be mapped to explicit, observable grading features - and the reliability of the human benchmark, rather than raw model capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。