arXiv:2606.11435cs.CL2026-06被引 5

系统梳理智能体技能评估与演化框架,推动从孤立技能到自动进化转变。

Agent Skill Evaluation and Evolution: Frameworks and Benchmarks

论文配图:Agent Skill Evaluation and Evolution: Frameworks and Benchmarks
图 1 · 摘自论文原文
  • 提出四类技能演化范式:执行反馈、轨迹蒸馏、压缩和强化学习。
  • 分析六类以技能为中心的基准,揭示覆盖盲区与指标不足问题。
  • 适合研究智能体系统、具身智能及自动化构建的学者与开发者。

智能体技能的增长重塑了智能体系统的构建、评估与部署方式。随着技能库规模持续扩大,严格的评估变得至关重要,以确保其在真实应用中的实用性、质量和安全性。因此,领域正经历从孤立技能生成向自动化、评估驱动的技能演化范式转变。本文系统考察了超越基础技能创建的技能演化与评估全景。我们将演化分为四大范式:执行反馈、轨迹蒸馏、压缩和强化学习,阐明各方法如何提升技能的实用性和可靠性。同时,我们分析了六类以技能为核心的基准,识别出基准覆盖结构上的缺口、权衡关系以及度量丰富性不足的问题,以推动技能研究进展。最后,我们指出了构建可泛化、高效且可验证安全的技能生态系统的开放方向。项目地址:https://github.com/Cassie07/AgentSkill_Survey。

原文摘要 · Abstract (English)

The growth of agent skills has transformed how agentic systems are built, evaluated, and deployed. As skill libraries continue to scale, rigorous evaluation becomes critical to ensuring their utility, quality, and safety in real-world applications. Consequently, the field is undergoing an emerging paradigm shift from isolated skill creation to automated, evaluation-driven skill evolution. In this survey, we systematically examine the landscape of skill evolution and evaluation beyond foundational skill creation. We categorize evolution into four distinct paradigms, spanning execution feedback, trajectory distillation, compression, and reinforcement learning, showing how each element contributes to improving skill utility and reliability. We also provide an analysis of six skill-centric benchmark categories, identifying structural gaps in benchmark coverage, trade-offs, and metric richness to advance skill research. Finally, we identify open directions for building skill ecosystems that are generalizable, efficient, and verifiably safe. The project URL is https://github.com/Cassie07/AgentSkill_Survey

智能体技能演化评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。