评测AI研究代理在真实任务中超越人类专家的能力
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- 构建7个开放性机器学习研究环境,对比人类与AI代理表现
- 2小时预算下,顶尖AI代理得分是人类的4倍;8小时后人类小幅领先
- 适合关注AI自主科研、智能体评估的研究者
前沿AI安全政策强调自动化AI研发(R&D)的重要性。然而,目前缺乏高现实性的评估,也无直接的人类性能对比。我们提出RE-Bench(Research Engineering Benchmark, v1),包含7个挑战性、开放式的机器学习研究工程环境,以及61名人类专家完成的71次8小时实验数据。结果显示,在8小时内,82%的人类尝试获得非零分,24%达到或超过我们的强基线解。通过最佳k采样和不同时间预算,我们将人类与多个公开前沿模型进行对比:在每环境总2小时预算下,最优AI代理得分达人类4倍;但在8小时预算下,人类略胜一筹,32小时时人类得分是顶级AI代理的2倍。定性分析表明,现代AI代理在多数ML领域具备深厚知识——例如,其编写的Triton内核比所有人类专家更快——且生成测试方案速度超人类十倍,成本极低。所有评估环境、人类数据、分析代码与代理轨迹均已开源。
原文摘要 · Abstract (English)
Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations for AI R&D capabilities, and none that are highly realistic and have a direct comparison to human performance. We introduce RE-Bench (Research Engineering Benchmark, v1), which consists of 7 challenging, open-ended ML research engineering environments and data from 71 8-hour attempts by 61 distinct human experts. We confirm that our experts make progress in the environments given 8 hours, with 82% of expert attempts achieving a non-zero score and 24% matching or exceeding our strong reference solutions. We compare humans to several public frontier models through best-of-k with varying time budgets and agent designs, and find that the best AI agents achieve a score 4x higher than human experts when both are given a total time budget of 2 hours per environment. However, humans currently display better returns to increasing time budgets, narrowly exceeding the top AI agent scores given an 8-hour budget, and achieving 2x the score of the top AI agent when both are given 32 total hours (across different attempts). Qualitatively, we find that modern AI agents possess significant expertise in many ML topics -- e.g. an agent wrote a faster custom Triton kernel than any of our human experts' -- and can generate and test solutions over ten times faster than humans, at much lower cost. We open-source the evaluation environments, human expert data, analysis code and agent trajectories to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。