arXiv:2409.11363cs.CLcs.AI2024-09被引 109

构建可复现性评估基准,测试AI能否自动复现科研论文结果

CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

  • 设计90篇论文的270个复现任务,覆盖三大学科
  • 最先进模型在最难任务上仅达21%准确率
  • 支持快速并行评估,助力研究者高效验证新代理

AI代理有望协助完成包括科学探究在内的多种重要任务。为推动有用代理的发展,需构建既具挑战性又贴近真实任务的基准。本文提出CORE-Bench(计算复现性代理基准),用于衡量代理在科学核心环节——计算复现性上的表现。该任务要求使用提供的代码和数据复现研究结果,是科学研究的基础。CORE-Bench包含90篇跨计算机科学、社会科学和医学领域的论文,共270个任务,涵盖三个难度层级及纯文本与图文任务。我们提供可快速并行评估的系统,相比串行实现节省数天时间。评估了两种基线代理:通用型AutoGPT和专用型CORE-Agent,分别基于GPT-4o和GPT-4o-mini。最佳表现仅为最难任务21%准确率,显示自动化科研任务仍有巨大提升空间。能复现已有工作的代理,是构建可开展新研究、验证改进其他代理性能的智能体的关键一步。我们期望CORE-Bench能提升科研可复现性,并推动未来研究代理发展。

原文摘要 · Abstract (English)

AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the development of useful agents, we need benchmarks that are challenging, but more crucially, directly correspond to real-world tasks of interest. This paper introduces such a benchmark, designed to measure the accuracy of AI agents in tackling a crucial yet surprisingly challenging aspect of scientific research: computational reproducibility. This task, fundamental to the scientific process, involves reproducing the results of a study using the provided code and data. We introduce CORE-Bench (Computational Reproducibility Agent Benchmark), a benchmark consisting of 270 tasks based on 90 scientific papers across three disciplines (computer science, social science, and medicine). Tasks in CORE-Bench consist of three difficulty levels and include both language-only and vision-language tasks. We provide an evaluation system to measure the accuracy of agents in a fast and parallelizable way, saving days of evaluation time for each run compared to a sequential implementation. We evaluated two baseline agents: the general-purpose AutoGPT and a task-specific agent called CORE-Agent. We tested both variants using two underlying language models: GPT-4o and GPT-4o-mini. The best agent achieved an accuracy of 21% on the hardest task, showing the vast scope for improvement in automating routine scientific tasks. Having agents that can reproduce existing work is a necessary step towards building agents that can conduct novel research and could verify and improve the performance of other research agents. We hope that CORE-Bench can improve the state of reproducibility and spur the development of future research agents.

AI代理可复现性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。