为不信任的智能体技能提供结构化安全审计与鲁棒增强
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills

- 将技能包跨文件整合为可复用能力单元,实现多文件协同安全审查
- 在404个包的测试中达到98.33%恶意风险召回率和98.89%攻击一致性
- 适合构建可信智能体生态系统的开发者与安全研究员使用
Agent Skills 将 SKILL.md 文件、脚本、参考文档和仓库上下文打包为可复用的能力单元,将预加载审计从单提示过滤转变为跨文件安全审查。现有防护机制常误报风险,且在语义保持的重写下难以识别恶意意图。本文将未授权智能体技能的预加载审计建模为鲁棒的三分类任务,提出 SkillGuard-Robust,结合角色感知证据提取、选择性语义验证与一致性保留判定。在 SkillGuardBench 及两个公开生态扩展上,通过五个大范围评估视图(254至404个包)验证。在404个包的预留集合上,精确匹配率达97.30%,恶意风险召回率为98.33%,攻击精确一致性达98.89%;在254个外部生态视图中,分别达到99.66%、100.00%和100.00%。结果表明:分解式包审计显著提升冻结与公开生态的鲁棒性,而更严苛的外部源迁移仍是开放挑战。
原文摘要 · Abstract (English)
Agent Skills package SKILL.md files, scripts, reference documents, and repository context into reusable capability units, turning pre-load auditing from single-prompt filtering into cross-file security review. Existing guardrails often flag risk but recover malicious intent inconsistently under semantics-preserving rewrites. This paper formulates pre-load auditing for untrusted Agent Skills as a robust three-way classification task and introduces SkillGuard-Robust, which combines role-aware evidence extraction, selective semantic verification, and consistency-preserving adjudication. We evaluate SkillGuard-Robust on SkillGuardBench and two public-ecosystem extensions through five large evaluation views ranging from 254 to 404 packages. On the 404-package held-out aggregate, SkillGuard-Robust reaches 97.30% overall exact match, 98.33% malicious-risk recall, and 98.89% attack exact consistency. On the 254-package external-ecosystem view, it reaches 99.66%, 100.00%, and 100.00%, respectively. These results support a bounded conclusion: factorized package auditing materially improves frozen and public-ecosystem robustness, while harsher external-source transfer remains an open challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。