arXiv:2606.05080cs.AIcs.LG2026-06被引 2

测试大模型能否持续优化复杂任务,发现坚持迭代比初始表现更重要。

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

论文配图:AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
图 1 · 摘自论文原文
  • 设计36个长期闭环优化任务,模拟真实科研工程流程。
  • 多数前沿模型在限时内进展有限,仅Claude Opus表现突出。
  • 强调持续反馈与时间感知对长期自主任务的关键作用。

科学与工程进步本质上是长期迭代过程:提出改进、实验验证、测量结果并持续优化。然而现有前沿模型基准多聚焦单轮响应或短周期代理轨迹,无法反映长时间跨度下的持续改进挑战。为此,我们提出AutoLab,一个面向超长时序闭环优化的新基准。AutoLab包含36个真实、专家精选的任务,涵盖系统优化、谜题挑战、模型开发和CUDA内核优化四大领域。每个任务以正确但故意低效的基线开始,要求代理在严格时间预算内完成改进。评估17个先进模型发现,成功关键不在于初始尝试质量,而在于持续进行基准测试、修改与实证反馈整合的能力。尽管Claude Opus-4.6表现出较强长周期优化能力,多数前沿模型(包括若干专有模型)仍提前终止或耗尽预算后进展甚微。结果凸显时间意识与持续迭代在自主代理中的重要性。我们开源完整基准、评估工具与任务资源,推动真正具备长时序能力代理的研究发展。

原文摘要 · Abstract (English)

Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifacts. Yet existing benchmarks for frontier models primarily evaluate either single-turn responses or short-horizon agent trajectories, failing to capture the challenges of sustained iterative improvement over extended time horizons. To address this gap, we introduce AutoLab, a new benchmark for ultra long-horizon closed-loop optimization. AutoLab consists of 36 realistic, expert-curated tasks spanning four diverse domains: system optimization, puzzle & challenge, model development, and CUDA kernel optimization. Each task begins with a correct but deliberately suboptimal baseline and challenges agents to improve it within a strict wall-clock budget. Evaluating 17 state-of-the-art models reveals the dominant predictor of success is not the quality of an agent's initial attempt, but its persistence in repeatedly benchmarking, editing, and incorporating empirical feedback. While claude-opus-4.6 exhibits strong long-horizon optimization capabilities, most frontier models, including several proprietary ones, either terminate prematurely or exhaust their budgets with minimal progress. These results underscore the importance of time awareness and persistent iteration in autonomous agents. We open-source the full benchmark, evaluation harness, and task artifacts, to accelerate research toward truly capable long-horizon agents.

长时序优化自主代理基准测试迭代改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。