提出技能持续学习新基准,测试智能体自主发现与进化技能的能力。
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

- 设计统一工作流框架,支持任务间技能迁移与演化。
- 长期学习使任务成功率最高提升8.43个百分点。
- 适合研究持续学习、技能演化与智能体自主性的学者。
随着自主智能体能力边界不断扩展,它们可通过即插即用的外部技能完成专业化任务。然而现有基准大多仅检验模型使用给定技能的能力,未评估其从经验中发现技能、失败后修复技能及长期维护技能库的能力。本文提出SkillFlow,一个包含20个任务族、共166项任务的基准,各任务族遵循领域无关执行流程(DAEF),定义统一的工作流框架,使任务可共享一致流程。智能体在代理式终身学习协议下进行评估:初始无技能,按族顺序解决任务,通过轨迹与评分驱动生成技能补丁并外化,携带更新后的技能库继续。实验显示显著能力差距:Claude Opus 4.6在终身学习下任务成功率从62.65%提升至71.08%(+8.43点);但高技能使用率未必带来高收益:Kimi K2.5虽使用66.87%技能,仅提升0.60点;Qwen-Coder-Next完成率仅44.58%,且相较基础设置出现退化。SkillFlow为该方向提供结构化测试平台,并深入分析技能发现、修补、迁移及其失败模式。
原文摘要 · Abstract (English)
As the capability frontier of autonomous agents continues to expand, they are increasingly able to complete specialized tasks through plug-and-play external skills. Yet current benchmarks mostly test whether models can use provided skills, leaving open whether they can discover skills from experience, repair them after failure, and maintain a coherent library over time. We introduce SkillFlow, a benchmark of 166 tasks across 20 families in which task construction within each family follows a Domain-Agnostic Execution Flow (DAEF) that defines an agent workflow framework, allowing these tasks to share a consistent workflow. Agents are evaluated under an Agentic Lifelong Learning protocol in which they begin without skills, solve tasks sequentially within each family, externalize lessons through trajectory- and rubric-driven skill patches, and carry the updated library forward. Experiments reveal a substantial capability gap. For Claude Opus 4.6, lifelong skill evolution improves task success from 62.65% to 71.08% (+8.43 points). However, high skill usage does not necessarily imply high utility: Kimi K2.5 gains only +0.60 points despite 66.87% skill usage, while Qwen-Coder-Next reaches only a 44.58% task completion rate and still regresses relative to the vanilla setting. SkillFlow contributes a structured testbed for this direction and an in-depth empirical analysis of skill discovery, patching, transfer, and their failure modes under lifelong evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。