arXiv:2506.22419cs.AIcs.CL2025-06NeurIPS被引 8

测试大模型复现论文改进的能力,发现顶尖模型仍难搞定已知优化。

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements

  • 用19个纳米GPT加速训练任务构建自动化评估基准
  • 即使给详细提示,顶尖模型也难以复现已有代码优化
  • 适合评估AI自主科研能力,尤其关注结果复现性

大型语言模型的快速发展有望推动科学进步。实现这一目标的关键能力是复现现有研究成果。为评估AI代理在活跃研究领域中复现结果的能力,我们引入了自动化大模型速通基准(Automated LLM Speedrunning Benchmark),基于社区参与的NanoGPT速通竞赛——在最短时间内训练GPT-2模型。每个19个速通任务提供先前记录的训练脚本,可选配三种提示格式:伪代码至论文风格的改进描述。记录执行快速,且改进涵盖从算法级到硬件感知的多样代码优化。这些特性使该基准对前沿的LLM训练优化问题既具可及性又真实。我们发现,结合当前最优推理模型与先进辅助框架的LLM,在给出详细提示时仍难以重现实已知的改进。因此,该基准提供了一种简单、非饱和的衡量指标,评估大模型自动化科学复现的能力——这是自主研究代理所必需但不充分的技能。

原文摘要 · Abstract (English)

Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce existing work. To evaluate the ability of AI agents to reproduce results in an active research area, we introduce the Automated LLM Speedrunning Benchmark, leveraging the research community contributions on the NanoGPT speedrun, a competition to train a GPT-2 model in the shortest time. Each of the 19 speedrun tasks provides the agent with the previous records training script, optionally paired with one of three hint formats, ranging from pseudocode to paper-like descriptions of the new records improvements. Records execute quickly by design and speedrun improvements encompass diverse code-level changes, ranging from high-level algorithmic advancements to hardware-aware optimizations. These features make the benchmark both accessible and realistic for the frontier problem of improving LLM training. We find that recent reasoning LLMs combined with SoTA scaffolds struggle to reimplement already-known innovations in our benchmark, even when given detailed hints. Our benchmark thus provides a simple, non-saturated measure of an LLMs ability to automate scientific reproduction, a necessary (but not sufficient) skill for an autonomous research agent.

大模型复现自动化评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。