arXiv:2605.11418cs.AIcs.CR2026-05被引 5

研究AI技能注册表的语义供应链攻击,揭示文本可操纵技能发现与选择。

Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry

论文配图:Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
图 1 · 摘自论文原文
  • 通过伪造SKILL.md文本操控技能检索、选择与审核流程。
  • 攻击使恶意技能在搜索中排名提升至80%的前十位置,选中率达77.6%。
  • 适用于关注AI代理安全与可信第三方集成的研究者和开发者。

自主AI代理通过模块化文件包扩展能力,其SKILL.md文件描述了使用时机与方式。这种设计虽支持按需扩展,但也引入语义供应链风险:自然语言元数据与指令可影响技能的发现、展示、选择与加载。我们针对三个注册阶段开展真实场景实验,使用实际ClawHub技能与合理注册机制。在发现阶段,短文本触发可操纵基于嵌入的检索,使攻击技能在配对测试中获得最高86%胜率,80%进入前十。在选择阶段,仅修改描述的框架即导致功能等价的恶意变体被选中,平均占比77.6%。在治理阶段,语义规避策略使恶意技能在36.5%至100%情况下逃避阻断判定。结果表明,SKILL.md并非静态文档,而是直接影响代理发现、信任与使用第三方能力的操作性文本。

原文摘要 · Abstract (English)

Autonomous AI agents increasingly extend their capabilities through Agent Skills: modular filesystem packages whose SKILL.md files describe when and how agents should use them. While this design enables scalable, on-demand capability expansion, it also introduces a semantic supply-chain risk in which natural-language metadata and instructions can affect which skills are admitted, surfaced, selected, and loaded. We study SKILL.md - only attacks across three registry-facing stages of the Agent Skill lifecycle, using real ClawHub skills and realistic registry mechanisms. In Discovery, short textual triggers can manipulate embedding-based retrieval and improve adversarial skill visibility, achieving up to 86% pairwise win rate and 80% Top-10 placement. In Selection, description-only framing biases agents toward functionally equivalent adversarial variants, which are selected in 77.6% of paired trials on average. In Governance, semantic evasion strategies cause malicious skills to avoid a blocking verdict in 36.5%-100% of cases. Overall, our results show that SKILL.md is not passive documentation but operational text that shapes which third-party capabilities agents find, trust, and use.

AI安全供应链攻击智能体语义操纵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。