高阶模型性能评估需更可靠的数据与裁判,否则结果失真。
Benchmarks Saturate When The Model Gets Smarter Than The Judge
- 人工修正数学数据集,去除噪声并标注问题类型
- 96.4%的裁判分歧中,原裁判错误,模型能力被误判
- 模型越强,对裁判质量要求越高,二者缺一不可
基准测试是追踪大语言模型进展的重要工具,但数据集和评估方法中的不准确性持续削弱其有效性。本文提出Omni-MATH-2,是对Omni-MATH数据集的人工修订版本,包含一个干净的精确答案子集(n=4181)和一个带标签的非标准子集(n=247)。每个问题均经审核,确保LaTeX可编译、可求解且可验证,包括补全缺失图表、标注需证明/估算/图像的问题,并清除冗余内容。该过程显著降低数据集引入的噪声,实现更精准的模型性能评估。通过对比GPT-5 mini与原始Omni-Judge在两子集上的表现,发现裁判间存在显著分歧。专家标注显示,在96.4%的裁判分歧中,Omni-Judge判断错误,表明其无法区分模型真实能力,即使在基准尚未饱和前已失效。随着题目难度提升,更可靠的裁判愈发关键,以避免裁判错误掩盖模型间真实差异。此外,两种裁判均未能识别标签子集中的当前失效模式,说明数据质量与裁判可靠性共同决定基准评估的准确性。
原文摘要 · Abstract (English)
Benchmarks are important tools to track progress in the development of Large Language Models (LLMs), yet inaccuracies in datasets and evaluation methods consistently undermine their effectiveness. Here, we present Omni-MATH-2, a manually revised version of the Omni-MATH dataset comprising a clean, exact-answer subset ($n{=}4181$) and a tagged, non-standard subset ($n{=}247$). Each problem was audited to ensure LaTeX compilability, solvability and verifiability, which involved adding missing figures or information, labeling problems requiring a proof, estimation or image, and removing clutter. This process significantly reduces dataset-induced noise, thereby providing a more precise assessment of model performance. The annotated dataset also allows us to evaluate judge-induced noise by comparing GPT-5 mini with the original Omni-Judge, revealing substantial discrepancies between judges on both the clean and tagged problem subsets. Expert annotations reveal that Omni-Judge is wrong in $96.4\%$ of the judge disagreements, indicating its inability to differentiate between models' abilities, even well before saturation of the benchmark occurs. As problems become more challenging, we find that increasingly competent judges become essential in order to prevent judge errors from masking genuine differences between models. Finally, neither judge identifies the present failure modes for the subset of tagged problems, demonstrating that dataset quality and judge reliability are both critical to develop accurate benchmarks of model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。