arXiv:2608.14905cs.CL2026-08

系统评估800条自主科研轨迹,发现当前AI代理失败主因是缺乏自我反思能力。

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

论文配图:How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
图 1 · 摘自论文原文
  • 构建覆盖科研全流程的100个真实前沿任务评估集
  • 识别出45种可复现的失败模式,均源于缺乏元认知循环
  • 揭示强模型仍会犯错,问题在模型本身而非流程设计

AI辅助科研正进入自动化新阶段,即从假设生成到论文发表的全链路自主研究(AutoResearch)。现有评估无法揭示代理如何运作或在何处失效:任务范围窄、仅衡量结果不分析过程、失败诊断缺乏系统性与细粒度。为此,我们提出AutoResearchEval,包含7个科学领域中100个基于真实前沿研究的任务,覆盖从构思、检索、执行、分析、写作到评审的完整科研生命周期。对8种模型-框架组合进行评估,生成800条自主研究轨迹,并进行过程级标注。基于这些数据,我们构建了45种实证驱动的失败模式分类体系(ARFT)。通过人类校准的代理评判管道,实现对完整轨迹与中间产物的细粒度归因。所有失败模式均指向同一核心缺陷:当前代理缺乏元认知循环——即无法将产出与发现比对,修正不一致,或质疑路径合理性。该缺陷在所有8组组合中普遍存在,包括最强模型,表明问题根植于模型本身,而非特定架构。本工作未测试流程干预能否弥补该缺陷。数据集与分类体系已公开,以推动自主科学发现研究。

原文摘要 · Abstract (English)

AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.

自主科研大模型评估失败分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。