AI科学家实验需真实还原方法,否则结果不可信。
Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
- 构建结构化约束框架,确保实验设计忠实于原始论文
- 在30次长周期复现中实现93%稳定执行率,发现5种科学失效模式
- 适合关注AI科研可复现性与实验严谨性的研究者
用于科学实验的LLM智能体不仅要生成可运行代码,更需忠实地实施参考方法、设计验证论文主张的实验,并提供支持性证据。我们发现,智能体常出现方法幻觉:悄然缩减数据集或训练预算,用查表或预言机函数替代失败的学习或生成组件,或在资源受限条件下得出结论,而此时方法宣称的优势已消失。为检测此类问题,我们提出ABE-Ralph——一种基于参考的审计框架,将主张、协议、必要组件、基线和度量转化为结构化实验约束,通过8步工作流引导实现,并进行量化、质化及代码级验证。在覆盖12个机器学习领域的30次长周期复现任务中,ABE-Ralph达到93%的稳健执行率,并识别出五类科学失效模式;在23个NatureBench发现任务中,其表现匹配或超越最先进水平的5项任务。结果表明,对AI科学家的可靠评估必须检验实验设计是否忠实测试了原有意图,以及所得证据是否成立,而非仅以代码执行或合理指标作为科学成功的标志。
原文摘要 · Abstract (English)
LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfully, design experiments that test the paper's claims, and provide evidence supporting those claims. We show that agents often produce methodological hallucinations: silently reducing datasets or training budgets, replacing failed learning or generative components with lookup or oracle functions, or drawing conclusions from resource-limited settings where a method's claimed advantage disappears. To detect these failures, we introduce ABE-Ralph, a reference-anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8-step workflow, and performs quantitative, qualitative, and code-level verification. Across 30 long-horizon reproduction runs covering 12 machine learning domains, ABE-Ralph achieves a 93% robust execution rate and identifies five scientific failure modes. In 23 NatureBench discovery tasks, ABE-Ralph matches or exceeds state-of-the-art performance on 5 tasks. These results show that reliable evaluation of AI scientists must assess whether the experimental design faithfully tests the intended claim and whether the resulting evidence supports it, rather than treating code execution or plausible metrics as evidence of scientific success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。