构建20个前沿科学任务集,评估大模型科研代理全流程能力
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
- 设计20个覆盖多领域的科研任务,涵盖从创意到迭代的完整研究流程
- 仅用前沿模型与简单架构,4项任务超越人类最先进水平,16项仍落后
- 开源全部任务与评测代码,推动自主科学智能发展
大型语言模型代理在推动科学研究方面具有巨大潜力。为加速这一进程,我们推出AIRS-Bench(人工智能科研基准),一套来自顶尖机器学习论文的20个任务。这些任务覆盖语言建模、数学、生物信息学和时间序列预测等多个领域。AIRS-Bench任务评估代理在整个科研生命周期中的能力——包括想法生成、实验分析与迭代优化——且不提供基准代码。该任务格式灵活,便于新增任务与不同代理框架的严格对比。我们使用前沿模型结合顺序与并行结构建立了基线。结果表明,代理在4项任务上超过人类最先进水平,但在其余16项任务中未能达到。即使在超越人类的任务中,代理也未触及底层任务的理论性能上限。这些发现表明AIRS-Bench尚未饱和,仍有巨大提升空间。我们已开源所有任务定义与评估代码,以促进自主科研智能的发展。
原文摘要 · Abstract (English)
LLM agents hold significant promise for advancing scientific research. To accelerate this progress, we introduce AIRS-Bench (the AI Research Science Benchmark), a suite of 20 tasks sourced from state-of-the-art machine learning papers. These tasks span diverse domains, including language modeling, mathematics, bioinformatics, and time series forecasting. AIRS-Bench tasks assess agentic capabilities over the full research lifecycle -- including idea generation, experiment analysis and iterative refinement -- without providing baseline code. The AIRS-Bench task format is versatile, enabling easy integration of new tasks and rigorous comparison across different agentic frameworks. We establish baselines using frontier models paired with both sequential and parallel scaffolds. Our results show that agents exceed human SOTA in four tasks but fail to match it in sixteen others. Even when agents surpass human benchmarks, they do not reach the theoretical performance ceiling for the underlying tasks. These findings indicate that AIRS-Bench is far from saturated and offers substantial room for improvement. We open-source the AIRS-Bench task definitions and evaluation code to catalyze further development in autonomous scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。