arXiv:2608.04975cs.SEcs.AI2026-08

修复科学编程评测缺陷后,大模型能力显著提升。

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

  • 对65个测试题逐题审计,发现263处缺陷,多数需专业背景才能识别。
  • 修正后模型子问题准确率从45%~60%升至84%~98%,主问题准确率从9%~27%升至69%~92%。
  • 新基准已公开,适合评估真正科学编程能力的模型与研究者使用。

SciCode是衡量语言模型科学编程能力的标准评测,包含需结合前沿科学理论与可运行代码的高阶问题。尽管其在政府与国家级实验室中广泛使用,近年性能却停滞不前:最强2026年模型子问题准确率稳定在60%左右,且新旧模型表现无差异。本文通过领域专家对全部65道题的逐题审计,发现263处缺陷,其中192处影响91%的主要题目,导致正确、遵循指令的解法被错误判定——原因包括不可复现的黄金答案、过严容差或自相矛盾的规范。这些缺陷中78%需物理或数学专业知识才能识别。我们修正所有可验证缺陷,仅补充良定义问题所需的必要说明,修复评分机制,收紧过松测试,每项修改均有理由并经第二位专家独立复核。在修正后的基准上重测12个前沿模型,结果显著改善:子问题准确率由45%~60%跃升至84%~98%,主问题准确率从9%~27%提升至69%~92%。这表明当前顶尖模型的科学编程能力远超原评测所反映水平——瓶颈不在模型本身,而在评测工具质量。我们发布带有完整审计轨迹的SciCode-Verified作为公共标准。

原文摘要 · Abstract (English)

SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code. It is a component of the Artificial Analysis Intelligence Index and a standing evaluation in government and national-laboratory suites. Yet its scores have recently plateaued: the strongest 2026 models cluster tightly around 60\% subproblem accuracy, and a successor model ties its predecessor. We trace this stagnation to defects in the benchmark itself. A per-problem, domain-expert audit of all 65 test problems uncovers 263 defects; 192 of them, spread across 91\% of the main problems, cause correct, instruction-following solutions to be wrongly rejected---through non-reproducible gold answers, over-tight tolerances, or self-contradictory specifications. Critically, 78\% of these score-suppressing defects require specialized physics or mathematics knowledge to detect, not mere clerical proofreading. We corrected every confirmable defect to produce SciCode-Verified. The corrections add only the specifications a well-posed problem requires, repair grading, and tighten the tests that were too lenient; every change is recorded with its justification and independently re-checked by a second domain expert. We re-evaluate twelve frontier model snapshots on the corrected benchmark and find a substantial recovery: subproblem accuracy rises from 45--60\% to 84--98\%, and main-problem accuracy from 9--27\% to 69--92\%. State-of-the-art models are far more proficient in scientific coding than SciCode has suggested---the bottleneck was not model capability, but the quality of the evaluation instrument. We release SciCode-Verified with its complete audit trail as the corrected public standard.

科学计算评测基准模型评估缺陷修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。