AI代理能独立做开放式科研吗?实测发现:能写代码,但难出真创新。
Can AI agents conduct open-ended AI research? Early evidence from two case studies

- 让AI代理复现顶会论文的核心研究问题,原作者评分评判
- 两次实验中,代理均未推进关键问题,被原作者明确拒绝
- 暴露出判断力差、创意不足、资源管理差等五大致命短板
AI自动化科研的前景依赖于智能体能否开展开放式研究。当前评估多局限于狭窄可验证任务或盲审生成论文,前者排除开放性,后者效率低且评审质量差。本文提出“影子评估”新方法:让代理承担高质量未发表论文的核心开放问题,由原作者评分。我们对两篇2026年NeurIPS投稿进行了测试,给予代理六天时间与数千美元算力。代理全程自主完成工程实现,但未能在核心问题上取得实质性进展,两位原作者均明确拒稿。识别出五类常见失败模式:对发表门槛判断失误、应对设计缺陷缺乏创意、死胡同回溯无效、资源意识薄弱、指令漂移。另一模型与框架的鲁棒性检验重现了这些失败。我们公开专家评审、问卷、代理仓库与日志。结果表明,当前代理可胜任科研工程,但在研究核心环节仍显著受限。
原文摘要 · Abstract (English)
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。