构建可复现性评估基准,测试AI能否自动复现科研论文结果
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
- 设计90篇论文的270个复现任务,覆盖三大学科
- 最先进模型在最难任务上仅达21%准确率
- 支持快速并行评估,助力研究者高效验证新代理
AI代理有望协助完成包括科学探究在内的多种重要任务。为推动有用代理的发展,需构建既具挑战性又贴近真实任务的基准。本文提出CORE-Bench(计算复现性代理基准),用于衡量代理在科学核心环节——计算复现性上的表现。该任务要求使用提供的代码和数据复现研究结果,是科学研究的基础。CORE-Bench包含90篇跨计算机科学、社会科学和医学领域的论文,共270个任务,涵盖三个难度层级及纯文本与图文任务。我们提供可快速并行评估的系统,相比串行实现节省数天时间。评估了两种基线代理:通用型AutoGPT和专用型CORE-Agent,分别基于GPT-4o和GPT-4o-mini。最佳表现仅为最难任务21%准确率,显示自动化科研任务仍有巨大提升空间。能复现已有工作的代理,是构建可开展新研究、验证改进其他代理性能的智能体的关键一步。我们期望CORE-Bench能提升科研可复现性,并推动未来研究代理发展。
原文摘要 · Abstract (English)
AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the development of useful agents, we need benchmarks that are challenging, but more crucially, directly correspond to real-world tasks of interest. This paper introduces such a benchmark, designed to measure the accuracy of AI agents in tackling a crucial yet surprisingly challenging aspect of scientific research: computational reproducibility. This task, fundamental to the scientific process, involves reproducing the results of a study using the provided code and data. We introduce CORE-Bench (Computational Reproducibility Agent Benchmark), a benchmark consisting of 270 tasks based on 90 scientific papers across three disciplines (computer science, social science, and medicine). Tasks in CORE-Bench consist of three difficulty levels and include both language-only and vision-language tasks. We provide an evaluation system to measure the accuracy of agents in a fast and parallelizable way, saving days of evaluation time for each run compared to a sequential implementation. We evaluated two baseline agents: the general-purpose AutoGPT and a task-specific agent called CORE-Agent. We tested both variants using two underlying language models: GPT-4o and GPT-4o-mini. The best agent achieved an accuracy of 21% on the hardest task, showing the vast scope for improvement in automating routine scientific tasks. Having agents that can reproduce existing work is a necessary step towards building agents that can conduct novel research and could verify and improve the performance of other research agents. We hope that CORE-Bench can improve the state of reproducibility and spur the development of future research agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。