用6种教育测评方法重评大模型,发现传统得分标准掩盖了真实能力差异。
Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

- 引入6种非纯正确数的评分方式,评估排除干扰项、自信度等能力
- 31个大模型排名因评分方式改变,且更符合用户实际偏好
- 揭示模型特性:GPT-5少回避,开放模型常犹豫不决
自然语言处理中的多选题评测普遍采用正确数量(准确率)评分,但在教育测评中,评分方案——即模型作答方式与评分规则的组合——是决定奖励何种能力的关键设计。本文考察六种受教育学启发的替代评分方案,用于评估超出准确率的能力,包括干扰项排除、放弃答题、置信度校准和自我修正。在大模型基准测试中,这些方案:1)使31个大模型的排名发生显著变化,超越仅改写问题的准确率提示;2)更准确预测用户在LLM Arena中的偏好;3)揭示不同模型的能力特征,例如GPT-5极少放弃答题且擅长自我修正,而较弱的开源模型则频繁回避且不愿排除选项。基于其优势,本文探讨将此类方案扩展至多选题以外任务的可能性。
原文摘要 · Abstract (English)
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。