arXiv:2608.18336cs.AIcs.CY2026-08

测试大模型在越南高考非线性评分下的真实表现,发现传统准确率会严重高估模型能力。

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

  • 构建真实评分规则的基准集THPT-Ladder,模拟越南高考第二部分的凸函数评分机制
  • 实测8个模型每题平均被扣0.020至0.159分,导致排名大幅下滑
  • 强调在非线性评分体系下,准确率无法反映真实水平,适合关注评估公平性的研究者

在评估语言模型于人类考试中的表现时,现有基准通常将每道题判为对或错,并报告整体准确率。这一方法假设部分知识应按比例计分,但在采用非加性评分规则的考试中该假设失效。2025年越南全国高中毕业考试改革即体现此问题:第二部分每题需判断四个真假陈述,得分呈凸函数分布——正确数为0、1、2、3、4时分别得0、0.10、0.25、0.50、1.00分。正确三题仅得0.50分,而非标准准确率所暗示的0.75分。因第二部分占总分10.00中的4.00分,使用准确率会夸大模型表现。本文引入THPT-Ladder基准,包含21份官方试卷中632个题目,严格按教育部评分标准执行。基于百万考生真实成绩,可直接将模型置于人群之中。在8个模型中,官方评分比比例计分少0.020至0.159分/题。该差距改变模型实际排名:以Qwen3.5-27B在2025历史考卷为例,0.042分差使其排名从第90百分位降至第77百分位(共481,293人)。模型准确率不能预测此惩罚;在与Claude Sonnet 5相同准确率下,不同错误分布使得分在0.869至0.932间波动。官方分数依赖正确项的组合方式,表明标准基准报告的能力是机构不会认可的虚假表现。

原文摘要 · Abstract (English)

When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.

评估基准非线性评分模型评测越南高考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。