分析2.6万项公开技能,发现中英文技能用途差异大,超30%存安全风险。
Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub

- 构建2.6万项技能数据集,对比中英文技能功能分布差异。
- 30%以上技能被标记为可疑或恶意,部分缺乏安全可观测性。
- 仅用发布时信息可预测风险,最佳模型准确率达72.62%。
技能生态系统已成为大型语言模型智能体系统的重要组成部分,支持任务封装、公开分发和社区驱动的能力共享。然而,尽管其快速发展,公共技能注册表的功能、生态结构及安全风险仍缺乏深入研究。本文对大规模公开技能注册表ClawHub开展实证研究,构建并标准化了包含26,502项技能的数据集,系统分析其语言分布、功能组织、流行度与安全信号。聚类结果显示显著的跨语言差异:英语技能更偏向基础设施,集中于API、自动化和记忆等技术能力;中文技能则更应用导向,形成媒体生成、社交内容生产、金融相关服务等场景化集群。进一步发现,超过30%的爬取技能被平台信号标记为可疑或恶意,且大量技能仍缺乏完整安全可观测性。为研究早期风险评估,我们基于发布时信息构建提交时间风险预测任务,建立包含11,010项技能的平衡基准。在12种分类器中,最优逻辑回归模型达到72.62%准确率与78.95% AUROC,主要文档成为最具信息量的提交时信号。研究揭示公共技能注册表既是智能体能力复用的关键推手,也是生态系统级安全风险的新暴露面。
原文摘要 · Abstract (English)
Skill ecosystems have emerged as an increasingly important layer in Large Language Model (LLM) agent systems, enabling reusable task packaging, public distribution, and community-driven capability sharing. However, despite their rapid growth, the functionality, ecosystem structure, and security risks of public skill registries remain underexplored. In this paper, we present an empirical study of ClawHub, a large public registry of agent skills. We build and normalize a dataset of 26,502 skills, and conduct a systematic analysis of their language distribution, functional organization, popularity, and security signals. Our clustering results show clear cross-lingual differences: English skills are more infrastructure-oriented and centered on technical capabilities such as APIs, automation, and memory, whereas Chinese skills are more application-oriented, with clearer scenario-driven clusters such as media generation, social content production, and finance-related services. We further find that more than 30% of all crawled skills are labeled as suspicious or malicious by available platform signals, while a substantial fraction of skills still lack complete safety observability. To study early risk assessment, we formulate submission-time skill risk prediction using only information available at publication time, and construct a balanced benchmark of 11,010 skills. Across 12 classifiers, the best Logistic Regression achieves a accuracy of 72.62% and an AUROC of 78.95%, with primary documentation emerging as the most informative submission-time signal. Our findings position public skill registries as both a key enabler of agent capability reuse and a new surface for ecosystem-scale security risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。