arXiv:2607.10113cs.AI2026-07中稿 · TMLR综述被引 2

系统梳理智能体技能库的演化生命周期,揭示其动态管理机制。

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries

论文配图:Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries
图 1 · 摘自论文原文
  • 构建六类技能分类与八阶段生命周期框架,统一研究视角。
  • 发现验证质量显著影响技能增强型强化学习效果。
  • 适合关注智能体长期演化、技能管理与评估标准的研究者。

大型语言模型智能体越来越多地将可复用的程序存储在模型之外,这些程序被称为‘技能’:包括代码函数、自然语言指令、SKILL.md 包、工作流图或可学习适配器等,供未来智能体检索和调用。本综述基于 124 篇 2023–2026 年论文的审计,将动态技能系统归纳为‘生命周期管理、验证、持续演化的实体仓库’。智能体通过交互收集证据,提出技能更新,经验证后准入,组织存储以支持检索与组合,修复或删除过时条目,并通过溯源与回滚机制管理共享。我们提出三个分析工具:第一,六类技能结构分类法;第二,八阶段生命周期架构,涵盖证据获取、提案、验证/准入、存储、检索/组合、维护、提炼/可移植性、治理;第三,轻量级技能记录格式与十种操作符词汇,实现跨文献比较。分析揭示:准入与修复反复关键;验证质量显著影响技能感知强化学习;扁平检索随库增长而退化;现有基准仍低估技能轨迹、使用-效用差距与安全边界。最后提出明确报告规范与开放问题,推动对动态技能库的评价从静态提示或工具集合转向演化系统视角。

原文摘要 · Abstract (English)

Large language model agents increasingly store reusable procedures outside the model. These reusable procedures are often called \emph{skills}: they may be code functions, natural-language instructions, SKILL.md packages, workflow graphs, or learned adapters that a future agent can retrieve and invoke. This taxonomy-driven survey asks how such skill libraries change over time. Across a $124$-paper $2023$--$2026$ audit set, we synthesize dynamic skill systems as \emph{lifecycle-managed, verified, evolving artifact stores}: agents collect evidence from interaction, propose skill updates, verify and admit candidates, organize them for retrieval and composition, repair or prune stale entries, and govern sharing through provenance and rollback. We organize the literature around three survey tools. First, a $\text{six}$-sense taxonomy distinguishes the structurally different artifacts called ``skills'' in current papers. Second, an $\text{eight}$-stage lifecycle architecture identifies the recurring design decisions behind evidence acquisition, proposal, verification/admission, storage, retrieval/composition, maintenance, distillation/portability, and governance. Third, a lightweight skill-record schema and $\text{ten}$-operator vocabulary provide common terms for comparing library updates without elevating them into a separate method contribution. Using this structure, we synthesize evidence-graded patterns with explicit caveats: admission and repair are repeatedly important, verifier quality materially affects skill-aware RL, flat retrieval can degrade as libraries grow, and current benchmarks still under-report library trajectories, usage--utility gaps, and safety surfaces. We close with concrete reporting standards and open problems for evaluating dynamic skills as changing libraries rather than static prompt or tool collections.

智能体技能库生命周期演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。