构建首个全面恶意技能检测基准,解决数据碎片化问题。
MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

- 整合13个来源,提炼7539个唯一恶意技能条目。
- 多模型测试显示检测准确率最高0.932,但跨源时骤降至0.665。
- 适合安全研究者和大模型防护系统开发者使用。
Agent Skills 为大模型代理提供可复用的指令包,但也成为恶意行为的直接传播渠道。现有恶意技能数据分散于多个来源,格式、证据标准与良性覆盖不一,且存在重复和结构关联内容,难以直接整合与评估。我们提出 MaliciousSkillBench,一个全面的恶意技能检测基准。整合13个公开来源,其中11个贡献核心恶意样本,将8,414条原始恶意记录归一化为7,539个唯一身份,分属4,588个操作结构家族。经交叉标签冲突剔除后,主基准包含9,740项技能:7,505项恶意,2,235项良性。通过统一11类攻击,对4,983个恶意身份进行来源映射,发现各来源威胁构成差异显著。我们评估三种文本检测器和三种现成技能扫描器。学习型检测器在随机划分下取得0.882-0.932的宏平均F1,但在源无关评估中降至0.653-0.665;最强的词TF-IDF SVM在随机/结构无关/源无关场景下分别得0.932/0.916/0.665,保持95.6%恶意召回率,但对未见来源产生62.4%的良性误报率。现成扫描器虽降低误报,但大幅牺牲恶意召回率。结果表明,可靠检测需兼顾跨源覆盖与检测与误伤的联合评估。
原文摘要 · Abstract (English)
Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration. This creates a direct distribution channel for malicious behavior, yet existing malicious-Skill datasets are fragmented across sources, artifact formats, evidence regimes, and benign coverage; duplicated and structurally related content further complicates direct aggregation and evaluation. We present MaliciousSkillBench, a comprehensive benchmark for malicious Agent Skill detection. We consolidate 13 public sources, 11 of which contribute Core malicious artifacts, and reduce 8,414 raw malicious records to 7,539 normalized-unique identities in 4,588 operational structural families. After conservative cross-label conflict exclusion, the primary benchmark contains 9,740 Skills: 7,505 malicious and 2,235 benign. To characterize its coverage, we harmonize 11 attack categories for 4,983 malicious identities with supported source-native mappings and find substantial differences in threat composition across sources. We then evaluate three learned text detectors and three off-the-shelf Skill scanners. Learned detectors achieve 0.882-0.932 Random Macro-F1 but only 0.653-0.665 under Source-Disjoint evaluation; the strongest word TF-IDF SVM scores 0.932/0.916/0.665 on Random/structural-disjoint/Source-Disjoint while retaining 95.6% malicious recall but producing 62.4% benign FPR on held-out sources. Off-the-shelf scanners occupy different but also unsatisfactory operating regimes, reducing false positives only at the cost of sharply lower malicious recall. Together, these results show that reliable malicious-Skill detection requires both broader cross-source benchmark coverage and evaluation that jointly measures attack detection and benign over-flagging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。