自动评估大模型技能生态,揭示真实场景下技能效果差异。
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

- 基于真实世界任务动态构建评测实例,替代静态基准
- 600+任务实例验证:技能并非越用越好,效果依赖模型与框架
- 适合研究者和开发者优化技能选型与部署策略
技能(Skills)即为大型语言模型(LLMs)提炼的结构化工作流指令,正成为提升智能体在现实任务中表现的关键机制。然而,随着开源技能生态快速扩张,不同模型与智能体框架如何与技能交互、如何评估技能质量、用户如何在实际成本与性能间权衡选择技能等问题仍不明确。本文提出 extsc{OpenSkillEval},一个面向技能增强型智能体系统及其自身能力的自动化评估框架。该框架不依赖静态基准,而是从五个下游应用领域(演示文稿生成、前端网页设计、海报生成、数据可视化、报告生成)的真实演化资源中自动构造出具有现实意义的任务实例。同时,收集并组织社区贡献的30个开源技能,在统一任务设置下进行可控比较。基于超过600个动态生成的任务实例,对主流模型与智能体框架进行系统评估。结果表明:技能可用性不等于有效使用;技能增强效果强烈依赖底层模型与智能体框架;许多公开流行的技能并未持续优于无技能基线。这些发现凸显了动态、任务驱动评估的重要性,并为技能的设计、选择与部署提供了实用洞见。更多案例与基准资源见项目网站:https://yingjiahao14.github.io/OpenSkillEval-Web/。
原文摘要 · Abstract (English)
Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improving agent performance on real-world downstream tasks. However, as the open-source skill ecosystem rapidly expands, it remains unclear how different models and agent frameworks interact with skills, how to evaluate skill quality, and how users should select skills under practical cost-performance trade-offs. In this paper, we present \textsc{OpenSkillEval}, an automatic evaluation framework for both skill-augmented agent systems and the skills themselves. Instead of relying on static benchmarks, \textsc{OpenSkillEval} automatically constructs realistic task instances from evolving real-world artifacts across five categories of downstream applications: presentation generation, front-end web design, poster generation, data visualization, and report generation. It further collects and organizes community-contributed skills for controlled comparison under unified task settings. Using more than 600 dynamically generated task instances and 30 open-source skills, we conduct a systematic evaluation of state-of-the-art models and agent frameworks. Our results show that skill availability does not guarantee effective skill usage, that the benefit of skill augmentation depends strongly on both the underlying model and the agent framework, and that many publicly popular skills do not consistently outperform base agents without skills. These findings highlight the need for dynamic, task-grounded evaluation and provide practical insights into the design, selection, and deployment of skills for LLM agents. Additional cases and benchmark resources are available on the project website: https://yingjiahao14.github.io/OpenSkillEval-Web/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。