arXiv:2604.17308cs.AI2026-04被引 17

提出技能持续学习新基准,测试智能体自主发现与进化技能的能力。

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

论文配图:SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
图 1 · 摘自论文原文
  • 设计统一工作流框架,支持任务间技能迁移与演化。
  • 长期学习使任务成功率最高提升8.43个百分点。
  • 适合研究持续学习、技能演化与智能体自主性的学者。

随着自主智能体能力边界不断扩展,它们可通过即插即用的外部技能完成专业化任务。然而现有基准大多仅检验模型使用给定技能的能力,未评估其从经验中发现技能、失败后修复技能及长期维护技能库的能力。本文提出SkillFlow,一个包含20个任务族、共166项任务的基准,各任务族遵循领域无关执行流程(DAEF),定义统一的工作流框架,使任务可共享一致流程。智能体在代理式终身学习协议下进行评估:初始无技能,按族顺序解决任务,通过轨迹与评分驱动生成技能补丁并外化,携带更新后的技能库继续。实验显示显著能力差距:Claude Opus 4.6在终身学习下任务成功率从62.65%提升至71.08%(+8.43点);但高技能使用率未必带来高收益:Kimi K2.5虽使用66.87%技能,仅提升0.60点;Qwen-Coder-Next完成率仅44.58%,且相较基础设置出现退化。SkillFlow为该方向提供结构化测试平台,并深入分析技能发现、修补、迁移及其失败模式。

原文摘要 · Abstract (English)

As the capability frontier of autonomous agents continues to expand, they are increasingly able to complete specialized tasks through plug-and-play external skills. Yet current benchmarks mostly test whether models can use provided skills, leaving open whether they can discover skills from experience, repair them after failure, and maintain a coherent library over time. We introduce SkillFlow, a benchmark of 166 tasks across 20 families in which task construction within each family follows a Domain-Agnostic Execution Flow (DAEF) that defines an agent workflow framework, allowing these tasks to share a consistent workflow. Agents are evaluated under an Agentic Lifelong Learning protocol in which they begin without skills, solve tasks sequentially within each family, externalize lessons through trajectory- and rubric-driven skill patches, and carry the updated library forward. Experiments reveal a substantial capability gap. For Claude Opus 4.6, lifelong skill evolution improves task success from 62.65% to 71.08% (+8.43 points). However, high skill usage does not necessarily imply high utility: Kimi K2.5 gains only +0.60 points despite 66.87% skill usage, while Qwen-Coder-Next reaches only a 44.58% task completion rate and still regresses relative to the vanilla setting. SkillFlow contributes a structured testbed for this direction and an in-depth empirical analysis of skill discovery, patching, transfer, and their failure modes under lifelong evaluation.

智能体终身学习技能演化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。