测试大模型代理如何持续学习并积累可复用技能。
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

- 构建动态评估框架,按难度排序100个任务,支持跨任务技能复用。
- 上下文学习效果接近显式技能维护,但提升主要来自反馈适应而非抽象复用。
- 能力较弱模型积累更多碎片化技能,难形成通用能力,适合研究技能演化。
现代代理框架为大语言模型配备外部技能库以解决复杂任务,但其能否有效进化技能及提升任务解决能力仍不明确。为此,我们提出ContinualSkillBench,一个用于上下文持续技能学习的动态评估框架。涵盖五个代表性领域,每个领域包含100个按难度递增排列且相互关联的子任务,支持跨任务技能复用。实验表明,顺序执行通常提升性能,但不同模型与领域间差异显著。上下文学习平均表现与显式技能维护相当,说明多数提升源于对先前上下文和反馈的适应,而非可复用技能的抽象。显式技能在需要可复用流程或精确输出的任务中仍有优势。此外,能力较弱模型倾向于积累更大、更碎片化的任务专属技能集合。结果表明,当前上下文技能演化机制可支持持续适应,但仍难以稳定地将经验凝练为鲁棒且可迁移的技能。
原文摘要 · Abstract (English)
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。