评估智能体在长期科研中的表现,发现其仍缺乏真正的创新研究能力。
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

- 用规则化指标分析智能体在任务中的方案设计、执行与反馈控制过程。
- 当前智能体多依赖已有技术组合,方法创新罕见,跨任务经验复用效果不稳定。
- 适合关注AI自主研发能力提升的研究者和系统设计者阅读。
自主智能体正逐步通过长期实验提升模型、系统及其他技术成果。然而,要真正理解其能力现状,评估必须超越最终得分——因为得分无法揭示进展或损失的具体环节,也无法判断累积经验是否有效提升后续决策。为此,我们基于新框架,对7个前沿模型在36个长期任务上进行了系统评估,采用规则化指标刻画任务内行为的解决方案构建、执行过程与反馈控制,并通过受控对比分析经验在任务内及跨任务间的复用情况。结果表明,当前智能体更像工程优化器而非完全自主的研究者:它们能提出并实施可行方案,但各次运行表现差异显著,最强解主要为已有技术的适应性调整或组合,真正的方法论创新极为罕见。详细分析显示,性能受多重因素影响,包括相似最终结果背后的差异化流程瓶颈、经验复用可能促进也可能误导后续决策,以及调用架构对性能稳定性的显著影响。这些发现为改进模型训练、推理策略、经验管理及调用设计提供了明确方向。
原文摘要 · Abstract (English)
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。