评测智能体技能在87个任务中的实际效果,发现精简技能提升显著。
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

- 构建87项任务的基准测试,对比有无技能时的表现差异。
- 使用技能后平均通过率从33.9%升至50.5%,提升16.6个百分点。
- 小模型加精简技能可媲美大模型,适合评估智能体能力优化。
智能体技能是推理时增强大型语言模型(LLM)智能体的结构化过程知识包。尽管迅速普及,目前尚无标准方法衡量其实际效用。我们提出SkillsBench,一个包含87项任务、覆盖8个领域,并配备精心筛选的技能与确定性验证器的基准。最新聚合评估在18种模型-执行配置下,对87项任务进行无技能与有技能条件的匹配测试。使用精选技能后,平均通过率从33.9%提升至50.5%(+16.6个百分点;归一化增益25.5%),配置级增益范围为+4.1至+25.7个百分点。聚焦型技能(最多三个模块)优于更庞大或全面的技能包,且小型模型加技能可达到大型模型无技能的表现。SkillsBench确立了配对评估作为衡量智能体技能在高专业性任务中有效性的重要基础。
原文摘要 · Abstract (English)
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。