arXiv:2606.10403cs.CL2026-06

用韩国高考数学题评估大模型推理能力,发现准确率骗人。

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

论文配图:KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty
图 1 · 摘自论文原文
  • 用十万级考生真实错题率构建数学推理基准
  • 发现模型在难题上表现差,且越算越错
  • 新指标能区分表面相似但本质不同的错误模式

数学推理评测已大量涌现,但多数缺乏基于真实人类表现的题目难度信号。我们提出KCSAT-ML,包含2014至2025年韩国大学学力水平考试(KCSAT;Suneung)数学题664道,其中339道核心题配有来自数十万考生的官方每题错误率。我们结合难度对齐推理增益(DRG)——一个与分数正交的度量,检验模型的错误是否集中在人类认为困难的题目上,还是容易的题目上。在多种视觉语言模型(VLMs)及通过OCR处理的大型语言模型(LLMs)中,我们发现三类模式:(i) 低预算下,所有规模模型在高人类错误率尾部性能骤降;(ii) 测试时缩放(TTS)使令牌消耗大致随考生错误率线性增长,而准确率提升呈非单调曲线;(iii) 同一模型家族内,TTS在最难题目上表现为反向缩放,在较易题目上则出现过度思考——二者皆为同一对齐失败的表现。在DRG指标上,准确率相近的模型可能处于截然相反的位置:一个模型错的是人类也难的题,另一个虽解出最难题却在人类易题上失分——这种差异被平均准确率掩盖。代码与数据构建工具将开源于https://github.com/naver-ai/KCSAT-ML。

原文摘要 · Abstract (English)

Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a decade (2014-2025) of Korean College Scholastic Ability Test (KCSAT; Suneung) mathematics: 664 problems with a 339-item core set carrying official per-item error rates from nationwide cohorts of hundreds of thousands of examinees. We pair the benchmark with Difficulty-aligned Reasoning Gain (DRG): a score-orthogonal metric that asks whether a model's mistakes concentrate on the items humans found hard, or on items humans found easy. Together they expose, across a wide range of VLMs (and LLMs via OCR), three patterns: (i) low-budget accuracy collapses on the high-human-error tail at every model size; (ii) test-time scaling (TTS) raises token use roughly linearly with cohort error rate, while accuracy gains follow a non-monotonic curve; (iii) within a single family, TTS flips between anti-scaling on the hardest items and overthinking on easier ones -- two faces of the same alignment failure. On DRG, models with near-identical accuracy can sit at near-opposite values: one model gets wrong what humans also find hard, while another solves the hardest items yet fails on items humans find easy -- a contrast that aggregate accuracy hides. Our code and dataset builder will be open-sourced at https://github.com/naver-ai/KCSAT-ML.

推理评估难度建模大模型测试KCSAT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。