用明确评分标准,小模型也能可靠批改开放题。
Grading Needs a Rubric, Not Intelligence

- 小模型按预设评分标准批改,不依赖自身智力。
- 答案本身解释了95.6%的分数差异,评卷人影响极小。
- 评分标准中的官方答案是关键,脱离它则评分变不可靠。
当基于明确评分标准时,小型语言模型可像昂贵模型一样可靠地批改开放题。我们以“任意到基准”(any-to-bench)为设计原则:前沿模型在数据摄入时一次性读取题目与评分标准;低成本模型完成重复批改任务。测试了两种模型族的六种配置,在三个推理强度水平下评估。每配置回答24道开放题,每份答卷被批改三次,共生成3,456次评分。分数主要取决于答案内容:答案身份解释了95.6%的分数方差,评卷人身份仅解释0.2%。提高写作者推理强度可使得分提升最多达全分值的0.143,而提高评卷人推理强度仅影响0.006。六个前沿级评卷人作为验证,结果一致且无更高可靠性。两个消融实验分解评分标准:移除评分标准和等级但保留官方答案,无显著影响;移除官方答案后,评分一致性从ICC 0.888降至0.628,分数虚高,评卷人推理再次产生影响。说明评分标准是解耦评分与评卷者能力的关键,而官方答案承担几乎全部判分工作。未发现长度偏好或同家族偏好。
原文摘要 · Abstract (English)
Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。