arXiv:2603.28815cs.CRcs.AI2026-03被引 7

评测智能体技能的实用性和安全性,助力构建更可靠的智能体系统。

SkillTester: Benchmarking Utility and Security of Agent Skills

  • 通过对比有无技能的执行结果,量化技能效用
  • 生成实用性得分、安全性得分及三级安全状态标签
  • 适合智能体开发者与安全研究人员使用

本技术报告介绍SkillTester,一个用于评估智能体技能实用性和安全性的工具。其评估框架结合成对的基线与带技能执行条件,以及独立的安全探测套件。基于比较实用性原则和用户友好性原则,框架将原始执行结果归一化为实用性得分、安全性得分和三级安全状态标签。总体而言,它可视为在智能体优先世界中对智能体技能进行比较质量保证的测试工具。公共服务已部署于https://skilltester.ai,项目代码托管于https://github.com/skilltester-ai/skilltester。

原文摘要 · Abstract (English)

This technical report presents SkillTester, a tool for evaluating the utility and security of agent skills. Its evaluation framework combines paired baseline and with-skill execution conditions with a separate security probe suite. Grounded in a comparative utility principle and a user-facing simplicity principle, the framework normalizes raw execution artifacts into a utility score, a security score, and a three-level security status label. More broadly, it can be understood as a comparative quality-assurance harness for agent skills in an agent-first world. The public service is deployed at https://skilltester.ai, and the broader project is maintained at https://github.com/skilltester-ai/skilltester.

智能体评估安全性工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。