arXiv:2607.05202cs.AI2026-07被引 7

评测智能体通过能力迁移实现自我进化,突破传统评估局限。

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

论文配图:EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
图 1 · 摘自论文原文
  • 基于执行轨迹提取可复用的能力单元,构建领域专属能力图谱。
  • 跨模型家族的能力迁移效果稳定,但现有自动方法无法全场景增益。
  • 适合研究智能体长期学习与自主进化机制的学者使用。

长周期大语言模型系统中的智能体自我进化主要表现为过程性:有用经验不仅是存储信息,更是可复用的搜索、调试与验证程序。然而当前评估方法未能分离这种形式的迁移。现有基准测试聚焦单次任务求解,记忆类基准则关注信息留存而非过程复用。我们提出EvoAgentBench,一个面向四类代理领域(网络调研、算法推理、软件工程、知识工作)的能力引导型自进化评测基准。该基准从代理执行中提取基于轨迹的能力,将其标准化为操作单元,并构建领域特定的能力图谱,连接具有过程重叠的任务。设计上,每个测试任务均有经过验证的训练侧能力支持。在528/267的训练/测试划分下,两种支架结构与三种骨干模型中,精心筛选的能力内容可在不同模型族间可靠迁移,但目前无自动方法能在所有设置中持续提升性能。EvoAgentBench将自进化评估从整体准确率比较转向对经验编码、路由与吸收的细粒度诊断。基准已公开于https://huggingface.co/datasets/EverMind-AI/EvoAgentBench。

原文摘要 · Abstract (English)

Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer. Agent benchmarks test single-episode task solving; memory benchmarks target information retention rather than procedural reuse. We introduce EvoAgentBench, a benchmark for agent self-evolution via Ability-guided transfer across four agentic domains: web research, algorithmic reasoning, software engineering, and knowledge work. EvoAgentBench extracts trace-grounded Abilities from agent executions, canonicalizes them into operational units, and builds domain-specific Ability Graphs linking tasks that share procedural overlap. By design, every test task is backed by verified training-side Ability support. Across a 528/267 train/test split, two scaffolds, and three backbones, curated Ability content transfers reliably across model families, but no current automatic method sustains positive gain in all settings. EvoAgentBench shifts self-evolution evaluation from aggregate accuracy comparison to fine-grained diagnosis of experience encoding, routing, and uptake. The benchmark is publicly available at https://huggingface.co/datasets/EverMind-AI/EvoAgentBench.

智能体自进化能力迁移评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。