分析3种工具对6.7万份AI技能的检测分歧,揭示代理安全需多层防护。
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree

- 构建包含67,453个技能的去敏数据集,整合三类扫描器结果
- 三类工具仅0.69%技能同时报警,81.9%仅单工具标记
- 技能安全需组合治理,不宜依赖单一工具判断
Agent skills 为 AI 代理提供可复用的指令、工具、脚本、参考和工作流,形成独立于模型安全与传统包级恶意软件检测的安全边界。ClawHub Security Signals 是一个包含 67,453 个最新公开 OpenClaw 技能版本的去敏数据集。每条记录包含脱敏后的 SKILL.md 内容及配套文件(如存在),并附有 ClawScan 注册表最终判定结果,以及来自三个扫描器家族的证据:VirusTotal、静态启发式分析和 NVIDIA SkillSpector。本文不估算恶意技能比例,而是研究扫描器之间的分歧。三类扫描器极少一致标记同一技能:任意两两重叠不超过其阳性总数的 10.4%,仅有 0.69% 的技能被三者共同标记,81.9% 的标记仅由单一扫描器产生。分歧呈现攻击面结构特征:SkillSpector 在 19,209 / 25,504 个可疑行中为阳性(75.3%),但仅在 14 / 206 个恶意行中为阳性(6.8%);而恶意判定区域则相反:206 个恶意行中有 150 个(72.8%)被 VirusTotal 标记,与捆绑代码的恶意证据一致。结果表明,代理技能安全需分层治理,而非依赖单一扫描器的允许/拒绝决策。该数据集作为去敏的银标准数据集发布:标签为注册表自动判定,非人工标注真值;发布版本为早期快照,旨在支持社区研究,同时等待人工标注子集开发。鼓励后续研究,包括针对技能安全筛查的专用模型。
原文摘要 · Abstract (English)
Agent skills extend AI agents with reusable instructions, tools, scripts, references, and workflows, establishing a security boundary distinct from both model safety and traditional package-malware detection. ClawHub Security Signals is a sanitized dataset of 67,453 latest public OpenClaw skill versions. Each row pairs redacted SKILL.md content and sanitized bundled files where present with a final ClawScan registry verdict and evidence from three scanner families: VirusTotal, static heuristic analysis, and NVIDIA SkillSpector. Rather than estimating malicious-skill prevalence, we study scanner disagreement. The three scanners rarely flag the same skills: any pair overlaps on at most 10.4% of their combined positives, only 0.69% of skills are flagged by all three, and 81.9% of flagged skills are identified by a single scanner. The disagreement is structured by attack surface. SkillSpector, which raises semantic agentic-risk advisories rather than malware-reputation signals, is positive for 19,209 of 25,504 suspicious rows (75.3%) but only 14 of 206 malicious rows (6.8%). The malicious-verdict region shows the inverse profile: 150 of 206 malicious rows (72.8%) are VirusTotal-positive, consistent with bundled-code malware evidence. These results show that agent-skill security requires layered governance, not single-scanner allow/block decisions. The corpus is released as a sanitized silver-standard dataset: labels are the registry's automated verdicts, not human-annotated ground truth, and the release represents an early, versioned snapshot intended to support the community while a human-annotated subset is developed. Further research is encouraged, including models tailored for skill-security triage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。