arXiv:2604.08988cs.AI2026-04被引 9

提出首个评估持续进化智能体的基准,揭示单一成功率会掩盖真实进化能力。

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

论文配图:SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
图 1 · 摘自论文原文
  • 基于飞轮理论构建可跨任务持续进化的智能体架构
  • 实测显示不同框架耗能差达31.2倍,且演化轨迹显著分化
  • 适合关注智能体长期学习与真实进化能力的研究者

当前基于大模型的智能体在单次任务中表现优异,但受限于静态工具集和记忆断层,无法跨任务积累经验。本文从数字具身视角定义自进化智能体(SEA),提出其最小充分架构——演进飞轮,并发布首个专为评估SEA设计的基准SEA-Eval。基于飞轮理论,该基准以成功率(SR)和任务耗能(T)为核心指标,通过连续任务流设计,量化演化增益、演化稳定性及隐式对齐收敛性。实证表明,在相似成功率达下,各框架个体任务耗能差异高达31.2倍,序列分析中演化轨迹明显分叉——说明仅看成功率会产生能力幻觉,而$T$的序列收敛才是区分真实进化与伪进化的关键标准。

原文摘要 · Abstract (English)

Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience across task boundaries. This paper formalizes the Self-Evolving Agent (SEA) from the perspective of digital embodiment and continuous cross-task evolution, introduces the Evolutionary Flywheel as its minimal sufficient architecture, and presents SEA-Eval -- the first benchmark designed specifically for evaluating SEAs. Grounded in Flywheel theory, SEA-Eval establishes SR and T as primary metrics and, through sequential task stream design, is designed to quantify evolutionary gain, evolutionary stability, and implicit alignment convergence. Empirical evaluation reveals that, under comparable success rates, token consumption differs by up to 31.2 times between frameworks on individual tasks, with divergent evolutionary trajectories emerging under sequential analysis -- demonstrating that success rate alone creates a capability illusion and that the sequential convergence of $T$ is the key criterion for distinguishing genuine evolution from pseudo-evolution.

智能体评估持续学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。