评测6个自动化研究系统在机器学习复现上的表现,发现其结果常有幻觉和漏洞。
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility

- 构建标准输入的全流程评测框架,用顶级论文重构实验任务。
- 仅10份生成论文通过自动评审,59%被接受的报告含虚构内容。
- 低成本系统比高算力系统更优,说明流程设计比算力更重要。
能够生成完整科学论文的自主研究系统发展迅速,但评估框架尚未跟上。为此,我们提出MLReplicate,一个端到端的机器学习复现性评测基准。该基准基于ICML 2025杰出论文重构为标准化输入,评估了6个前沿研究系统(AI SCIENTIST-V1、AI SCIENTIST-V2、AGENT LABORATORY、CYCLERESEARCHER、AI RESEARCHER、TINY SCIENTIST),共生成45篇论文,其中3次实验失败。评测采用双协议方法:自动化会议评审与结构化专家人工评估,并记录计算成本、运行时间与人工干预量。自动化评审中,37份有效提交中有10份被接受,另有8份因未达最低页数要求被直接拒稿。人工评审则一致发现各系统存在方法缺陷、实验结果幻觉及复现失败问题,59%通过自动评审的论文包含伪造或无依据主张。进一步发现,输入令牌量与计算成本无法预测输出质量:最廉价系统在人工评估中优于最耗能系统,尽管前者输入令牌仅为后者的1/38。这表明自主研究工作流设计的重要性超过算力规模。MLReplicate揭示了当前系统与真实科学严谨性的巨大差距,并建立了一个可扩展、实用的评估框架,推动可信的AI驱动科研发展。
原文摘要 · Abstract (English)
Autonomous research systems capable of generating complete scientific manuscripts have advanced rapidly, yet robust and realistic evaluation frameworks have failed to keep pace. To bridge this gap, we introduce MLReplicate, an end-to-end benchmark evaluating autonomous research systems on machine learning reproducibility. The benchmark was constructed from ICML 2025 outstanding papers reformulated into standardized input specifications and evaluated across 6 state-of-the-art research systems: AI SCIENTIST-V1, AI SCIENTIST-V2, AGENT LABORATORY, CYCLERESEARCHER, AI RESEARCHER, and TINY SCIENTIST, yielding 45 generated manuscripts, with 3 failed experiments. Outputs are assessed using a dual-protocol approach that combines automated conference-style review and structured expert human evaluation, while tracking computational cost, runtime, and the amount of required human intervention. The automated conference-style review accepted 10 out of 37 valid submissions. An additional 8 submissions were desk-rejected before review for failing to meet the minimum page threshold. In contrast to automated reviews, human reviewers consistently identified methodological flaws, hallucinated experimental results, and reproducibility failures across all systems, and 59% of accepted automated reviews contained fabricated or unsupported claims. We further find that neither token budget nor computational cost predicts output quality: the cheapest system outperforms the most resource-intensive system in human evaluation, despite a 38-fold difference in input tokens. We thus demonstrate that autonomous research workflow design matters more than the scale of compute. MLReplicate exposes a substantial gap between current autonomous research systems and genuine scientific rigor, and establishes a practical, extensible evaluation framework for systematic progress toward trustworthy AI-driven scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。