用AI自动复现论文,比人类更准更稳。
Training AI Scientists to Replicate Research

- 构建可扩展的论文复现任务空间,用自动生成评分标准提供奖励信号。
- 270亿参数的AI科学家Faraday在复现任务上超越Claude Opus 4.8和GPT-5.5。
- AI复现过程更符合科学思维,适合科研自动化与可信研究验证。
论文的可复现性是科学知识的基石,确保已有结果的可靠性并为后续实验提供基础。复现过程通常揭示此前未明确定义的细节,因此需要类似假设驱动的探索,而非开放式的科研。本文开发了Replica——一个可扩展的论文复现任务空间。为提供奖励信号,我们引入基于自动生成评分标准的判别器,噪声低且与人类对复现质量的评估高度一致。我们对一款270亿参数的“AI科学家”代理Faraday进行后训练,使其利用编码工具完成任务,在保留测试任务上表现优于Claude Opus 4.8和GPT-5.5。对单次运行的定性分析显示,Faraday采用更具科学原则性的方法。我们认为,这些结果为无需复杂框架的长程科学创新类AI代理提供了重要起点。
原文摘要 · Abstract (English)
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter "AI Scientist" agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。