测试AI能否复现天体物理论文,发现当前模型成功率不足20%。
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
- 构建可复现的天体物理论文任务框架,涵盖实验设置与推导分析
- 顶尖语言模型在复现任务中准确率低于20%,暴露严重可靠性问题
- 适合关注AI科研助手可信度的研究者与开发者参考
前沿AI代理在科学研究辅助方面展现出日益增长的潜力,但要用于原创性研究,必须先评估其工作的真实性和正确性。为此,我们提出ReplicationBench,一个评估代理是否能复现天体物理文献中完整论文的框架。天体物理依赖档案数据与计算研究,极少需实地实验,是测试AI科研助手的理想场景。每篇论文被拆分为多个任务,要求代理复现核心贡献,包括实验设置、推导过程、数据分析和代码实现,所有任务均与原作者共同设计,针对关键科学结果,实现对忠实度(遵循原方法)与正确性(技术准确性)的客观评估。该基准对当前前沿语言模型极具挑战:即使表现最佳的模型,整体得分也低于20%。通过与领域专家协作分析轨迹,我们揭示了代理在科学研究中多样且复杂的失败模式。ReplicationBench建立了首个基于论文级别的、经专家验证的天体物理研究任务基准,揭示了可推广至其他数据驱动科学领域的代理性能洞见,并提供了一个可扩展的框架,用于衡量AI代理在科研中的可靠性。
原文摘要 · Abstract (English)
Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use agents for novel research, we must first assess the underlying faithfulness and correctness of their work. To evaluate agents as research assistants, we introduce ReplicationBench, an evaluation framework that tests whether agents can replicate entire research papers drawn from the astrophysics literature. Astrophysics, where research relies heavily on archival data and computational study while requiring little real-world experimentation, is a particularly useful testbed for AI agents in scientific research. We split each paper into tasks which require agents to replicate the paper's core contributions, including the experimental setup, derivations, data analysis, and codebase. Each task is co-developed with the original paper authors and targets a key scientific result, enabling objective evaluation of both faithfulness (adherence to original methods) and correctness (technical accuracy of results). ReplicationBench is extremely challenging for current frontier language models: even the best-performing language models score under 20%. We analyze ReplicationBench trajectories in collaboration with domain experts and find a rich, diverse set of failure modes for agents in scientific research. ReplicationBench establishes the first benchmark of paper-scale, expert-validated astrophysics research tasks, reveals insights about agent performance generalizable to other domains of data-driven science, and provides a scalable framework for measuring AI agents' reliability in scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。