arXiv:2606.17819cs.SEcs.AI2026-06被引 4

为大模型智能体技能提供可扩展的评估框架,量化其实际效用。

A Framework for Evaluating Agentic Skills at Scale

论文配图:A Framework for Evaluating Agentic Skills at Scale
图 1 · 摘自论文原文
  • 构建真实任务评估技能关键能力,支持技能作者自定义验证。
  • 在500个真实技能上生成1000个任务,发现模型表现差异显著。
  • 揭示技能能改变模型行为,适合需定制化工作流的研究者使用。

智能体技能——可复用的知识构件,用于增强大语言模型智能体能力——已在工业界快速采用,但其跨领域影响及在商业与开源模型中的应用仍缺乏研究,且尚无通用的技能评估方法。本文提出一种评估框架,使技能作者可构建真实任务,严格评估其关心的技能特性,并通过完成任务来估算技能效用。我们基于500个真实技能,生成1000个任务,并制定指令遵循与目标完成评分标准。利用这些指标,评估了19种代理模型配置(含专有和开源)的表现。结果表明,不同模型在遵循技能内嵌指令上的差异导致性能提升显著不同。此外,相较于无技能场景,引入技能会显著改变模型行为,成为向大模型智能体编码特定工作流的关键机制。我们公开了该评估数据集,以支持未来智能体技能研究。

原文摘要 · Abstract (English)

Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet their cross-domain impact and use across commercial and open-source models remain under-studied, and no reusable methodology exists for evaluating an individual skill. In this work, we present an evaluation framework that lets a skill author construct realistic tasks to rigorously assess the aspects of a skill that matter most to them, and that estimates skill utility by solving those tasks. Further, we apply our evaluation approach at scale to 500 real-world skills, generating 1,000 tasks derived from the skills' content, along with instruction-following and goal-completion scoring rubrics. Using these metrics, we evaluate how 19 agent-model configurations, both proprietary and open-source, perform on the tasks. Our results show that models vary widely in how closely they adhere to the instructions encoded in skills, leading to substantial differences in their performance gains. Furthermore, we show that access to a skill significantly changes model behavior compared to the no-skill setup, providing an essential mechanism for encoding opinionated workflows into LLM agents. We release our evaluation dataset to support future work on agent skills.

智能体评估大模型技能可扩展评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。