让大模型自动创建、记忆和优化技能,持续提升任务解决能力。
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

- 构建技能全生命周期管理框架,实现技能自动生成与迭代。
- 在SkillsBench上自创技能达85.24%成功率,超人工技能(81.17%)。
- 技能可跨模型迁移,对Hermes迁移准确率达51.90%,适合长期进化系统研究。
大型语言模型(LLM)代理依赖可复用技能完成复杂任务,但现有技能生成方法常将技能视为孤立、静态的产物,限制了复用性、可靠性与长期改进。本文提出MUSE-Autoskill Agent(Memory-Utilizing Skill Evolution),一个以技能为中心的代理框架,实现技能的统一生命周期管理:创建、记忆、管理、评估与优化。MUSE按需生成技能,跨任务存储,并通过技能目录检索;同时积累每项技能的使用经验,用于后续复用与适应。在SkillsBench和SkillLearnBench的主要设置下,MUSE-Autoskill超越Hermes、Codex和Claude Code。在SkillsBench上,其自创技能在成功覆盖子集上达到85.24%成功率,超过人工编写的81.17%;且其生成的技能在迁移至Hermes时表现优于Codex或Claude生成的技能,准确率达51.90%。结果表明,将技能视为具有长期记忆、经验感知与可测试性的资产至关重要。
原文摘要 · Abstract (English)
Large language model (LLM) agents rely on reusable skills to solve complex tasks, but existing skill creation approaches often treat skills as isolated, static artifacts, limiting reusability, reliability, and long-term improvement. We propose MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution), a skill-centric agent framework that creates, reuses, and refines skills under a unified lifecycle: creation, memory, management, evaluation, and refinement. MUSE creates skills on demand, stores them across tasks, retrieves them through a skill catalog, and accumulates per-skill experience for later reuse and adaptation. Across the main reported settings on SkillsBench and SkillLearnBench, MUSE-Autoskill outperforms Hermes, Codex, and Claude Code. On SkillsBench, its self-created skills surpass human-authored skills on the successfully covered subset (85.24% vs. 81.17%), showing that lifecycle-managed skills can distill agent experience into highly effective reusable assets; MUSE-created skills also transfer to Hermes more effectively than Codex- or Claude-created skills, reaching 51.90% accuracy under transfer. These results highlight the importance of treating skills as long-lived, experience-aware, and testable assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。