arXiv:2512.11183cs.LG2025-12

用可衡量的科学进展替代静态题目,让模型评估真正推动研究前进。

Progress over Points: Reframing LM Benchmarks Around Scientific Objectives

  • 构建动态实验环境,以科学突破为评测目标而非固定题库
  • 训练速度提升3秒,达成新纪录并观察到新算法思路涌现
  • 适合关注模型训练效率与科研创新的研究者

当前主流语言模型评估依赖静态、已解决的问题(如数学应用题),虽能验证基础能力,但导致评测体系趋于僵化。为此,我们提出面向进展的评测范式:将评测目标设为科学进步本身,使模型性能提升直接推动领域发展。作为初步尝试,我们基于NanoGPT速度跑挑战构建了一个标准化环境,包含数据子集、参考模型、训练框架和实时验证机制,并设有防作弊检查。评估聚焦于科学增量——最优损失值与效率前沿。实验中,我们实现训练时间新纪录,较前人快3秒,并观察到新型算法思想的出现。模型间对比仍可进行,但仅为手段而非目的;核心目标是催生可复用的语言建模技术改进。本工作旨在推动社区从静态排行榜转向开放且可度量的测试期科学研究,使‘评测’成为科学进步的驱动力。

原文摘要 · Abstract (English)

Current benchmarks that test LLMs on static, already-solved problems (e.g., math word problems) effectively demonstrated basic capability acquisition. The natural progression has been toward larger, more comprehensive and challenging collections of static problems, an approach that inadvertently constrains the kinds of advances we can measure and incentivize. To address this limitation, we argue for progress-oriented benchmarks, problem environments whose objectives are themselves the core targets of scientific progress, so that achieving state of the art on the benchmark advances the field. As a introductory step, we instantiate an environment based on the NanoGPT speedrun. The environment standardizes a dataset slice, a reference model and training harness, and rich telemetry, with run-time verification and anti-gaming checks. Evaluation centers on the scientific delta achieved: best-attained loss and the efficiency frontier. Using this environment, we achieve a new state-of-the-art training time, improving upon the previous record by 3 seconds, and qualitatively observe the emergence of novel algorithmic ideas. Moreover, comparisons between models and agents remain possible, but they are a means, not the end; the benchmark's purpose is to catalyze reusable improvements to the language modeling stack. With this release, the overarching goal is to seed a community shift from static problem leaderboards to test-time research on open-ended yet measurable scientific problems. In this new paradigm, progress on the benchmark is progress on the science, thus reframing "benchmarking" as a vehicle for scientific advancement.

语言模型科学评测训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。