首个评估大模型自进化任务框架能力的基准,发现其在搜索任务中表现超人工设计。
Evo-Bench: Can Language Models Improve Agent Harness?

- 用辅助任务演化筛选敏感任务,分层拆分保证泛化性
- 顶尖模型在通用任务上提升16.6分,接近人工最优水平
- 自演化框架可迁移,但对复杂办公流程适应差
大型语言模型推动了自主代理的快速发展,但现有评估仍局限于静态任务求解。新兴前沿是“框架演化”——代理自主优化自身操作框架的能力。然而,系统性评测这一能力仍具挑战:现有方法无法区分框架改进与模型基础性能、易导致任务特异性过拟合、难以捕捉长周期迭代研究。为此,我们提出Evo-Bench,首个面向搜索、办公与通用代理领域的框架演化能力基准。通过新颖的框架引导构建范式,利用辅助任务演化识别对框架改进敏感的任务,并采用敏感性感知分层拆分确保跨套件鲁棒泛化。对九个前沿及开源模型的广泛评估显示,顶级模型绝对提升达16.6分,接近最先进的人工工程基线。关键发现:自主演化在通用任务中优于人工框架,在搜索任务中表现卓越,但在需高度特定处理流程的办公任务中表现受限。分析还揭示早期饱和等时间异常现象,且合成框架可作为高可迁移推理结构,持续提升多种策略模型性能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。