arXiv:2606.22613cs.AI2026-06被引 1

为大模型技能提供可自动评估的综合审查框架

SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment

论文配图:SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment
图 1 · 摘自论文原文
  • 基于技能包自动生成适配的任务进行评估
  • 发现超7%真实技能存在风险状态
  • 适合技能开发者与平台审核者使用

智能体技能已成为扩展大语言模型代理的有效方式,但其生态系统仍缺乏可靠的评估机制。现有方法多依赖固定任务集,在预设任务和环境中评估技能表现,难以区分技能本身贡献与基础模型能力,且易忽略技能在非预期任务中的价值。本文提出 SkillAudit,一种面向技能中心的端到端评估框架,输入任意代理技能后可自动生成涵盖实用性、效率/成本、安全性的多维度评估报告。该框架聚焦技能本身,从技能包中直接构建能力对齐的任务,并在隔离沙箱中执行以收集证据,再通过大模型判断生成可审计结果。为衡量实用性和效率,提出基线对比原则;为评估安全风险,设计静态语义分析与动态运行时验证相结合的两阶段检测机制。对涵盖23个职业类别的23个高排名真实技能包扫描后发现,超过7%的技能处于风险状态。

原文摘要 · Abstract (English)

Agent skills have become a practical way to extend large language model agents, but the growing skill ecosystem still lacks a reliable way to judge whether a skill is worth deploying. Existing evaluation methods remain largely anchored to fixed task suites, assessing skills through performance on predefined tasks and environments. As skill marketplaces expand, this paradigm becomes inadequate: fixed suites can conflate a skill's marginal contribution with backbone strength and miss its value when tasks fall outside the skill's intended scope. We introduce SkillAudit, an end-to-end framework for skill-centered assessment that takes an arbitrary agent skill as input and automatically generates a comprehensive, multi-dimensional evaluation report spanning utility, efficiency/cost, and safety. SkillAudit focuses on the skill artifact itself and constructs capability-aligned evaluation tasks directly from the skill package. The generated tasks are conducted in isolated sandbox environments to collect execution evidence, followed by automated checks with LLM-based judging to produce auditable results. To dissect the agent skills, we propose the baseline comparison principle to measure utility and efficiency/cost, and introduce a two-stage detection paradigm combining static semantic analysis with dynamic runtime verification to assess safety risks. After scanning top-ranked real-world skill packages spanning 23 occupational categories, we found that over 7% of skills are at risky status.

智能体评估技能审计安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。