AI科学家能写论文,但难验证真伪,这篇综述揭了科研自动化中的可信度缺口。
Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap

- 系统性筛查125篇论文,聚焦代码与结论的可验证性差距
- 仅38%发布实验种子或执行轨迹,无一例外部验证的闭环智能体
- 为审稿人提供可操作的报告清单,推动可复现研究标准
大型语言模型(LLM)代理正被广泛应用于科学研发全周期:从选题、文献检索、实验设计与执行、分析到论文撰写与评审。端到端的AI科学家系统已能生成类似论文的文本,但其声明往往比代码更难验证。本综述聚焦计算人工智能/机器学习研究中的这一验证鸿沟,涵盖代码、基准测试、实验和文稿等最可见环节。我们筛选125项候选工作,最终纳入35篇,对其中26篇进行全文编码:24个可运行系统,2篇研究或立场论文。编码七项审计维度:研发阶段、自主程度、评估方法、发布的成果物、人工介入点、新颖性验证方式、结果选择披露。主要发现是:代码发布已普遍(24个系统中83%),但具备可复现级别的数据与可验证声明的成果仍极少见;仅38%公开实验种子或执行轨迹,38%报告了新颖性验证方法。在9个闭合环路的L4级系统中,7个为机械重复,1个为作者自证且无外部检验;在所有收录的LLM时代系统中,均无符合编码规则的外部验证的内部奥义(in-loop oracle)。本文贡献包括一个编码语料库、生命周期-自主度映射图、可审计性缺口分析,以及面向审稿人的报告检查表。综述指出,当前领域核心瓶颈已不仅是代理能否完成科研任务,更是审稿人能否验证其所产声明的真实性。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review. End-to-end AI scientist systems can now produce paper-like manuscripts, but their claims are often harder to verify than their code is to run. This survey studies that gap in computational AI/ML research, where code, benchmarks, experiments, and write-ups are most visible. We screen 125 candidate works and include 35, with full-text coding of 26 entries: 24 runnable systems and two study or position papers. We code seven audit dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty verification, and result-selection disclosure. The main pattern is that code release is now common, but reproducibility-grade and claim-verification artifacts remain much less common. In the 24 runnable systems, 83 percent release code, while 38 percent release seeds or execution traces and 38 percent report any novelty-verification method. Among nine closed-loop L4 systems, seven are mechanical reruns and one is author-claimed without an external check; no LLM-era system in the corpus demonstrates an externally validated in-loop oracle under our coding rule. We contribute a coded corpus, a lifecycle-by-autonomy map, an auditability-gap analysis, and a reviewer-facing reporting checklist. The survey argues that the field's central bottleneck is no longer only whether agents can complete research tasks, but whether reviewers can verify the claims those agents produce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。