arXiv:2604.15415cs.CRcs.AI2026-04被引 7

首次系统检测智能体技能中的有害行为,发现近5%技能具潜在危害。

HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

论文配图:HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
图 1 · 摘自论文原文
  • 基于自建分类体系,用大模型评分识别有害技能
  • 98,440个技能中4.93%有害,爪子库有害率高达8.84%
  • 构建首个真实场景下的安全评估基准,助力模型防护研究

大型语言模型已演变为依赖开放技能生态(如ClawHub和Skills.Rest)的自主智能体,这些生态提供大量可公开复用的技能。现有安全研究多聚焦于技能内部漏洞(如提示注入),但对可能被滥用进行有害行为(如网络攻击、欺诈、隐私侵犯、色情内容生成)的技能——即有害技能——缺乏系统性评估。本文首次开展大规模测量研究,覆盖两大主流注册表中的98,440个技能。基于自建有害技能分类体系,通过大模型驱动的评分系统发现:4.93%的技能(4,858个)具有有害性,其中ClawHub的有害率为8.84%,远高于Skills.Rest的3.49%。进一步构建HarmfulSkillBench,首个在真实智能体情境下评估安全性的基准,包含200个跨20类别的有害技能及四种评估条件。对六款LLM的测试显示,通过预安装技能呈现有害任务,显著降低模型拒绝率:无技能时平均危害得分0.27,有技能时升至0.47,若意图隐含则达0.76。研究已向相关注册表负责任披露,并开源基准以支持后续研究(见https://github.com/TrustAIRLab/HarmfulSkillBench)。

原文摘要 · Abstract (English)

Large language models (LLMs) have evolved into autonomous agents that rely on open skill ecosystems (e.g., ClawHub and Skills.Rest), hosting numerous publicly reusable skills. Existing security research on these ecosystems mainly focuses on vulnerabilities within skills, such as prompt injection. However, there is a critical gap regarding skills that may be misused for harmful actions (e.g., cyber attacks, fraud and scams, privacy violations, and sexual content generation), namely harmful skills. In this paper, we present the first large-scale measurement study of harmful skills in agent ecosystems, covering 98,440 skills across two major registries. Using an LLM-driven scoring system grounded in our harmful skill taxonomy, we find that 4.93% of skills (4,858) are harmful, with ClawHub exhibiting an 8.84% harmful rate compared to 3.49% on Skills.Rest. We then construct HarmfulSkillBench, the first benchmark for evaluating agent safety against harmful skills in realistic agent contexts, comprising 200 harmful skills across 20 categories and four evaluation conditions. By evaluating six LLMs on HarmfulSkillBench, we find that presenting a harmful task through a pre-installed skill substantially lowers refusal rates across all models, with the average harm score rising from 0.27 without the skill to 0.47 with it, and further to 0.76 when the harmful intent is implicit rather than stated as an explicit user request. We responsibly disclose our findings to the affected registries and release our benchmark to support future research (see https://github.com/TrustAIRLab/HarmfulSkillBench).

智能体安全有害技能大模型风险基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。