提出可自动化构造的技能攻击基准,揭示代理系统在全生命周期中的潜在风险。
SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

- 构建自动化的攻击生成流水线,支持固定投毒与自我变异投毒两种场景。
- 覆盖71个技能、879个攻击样本,攻击成功率最高达86.3%。
- 揭示当前防御机制失效的根本原因:代理未调用污染文件而非真正抵抗。
智能体技能在工作流中处于特权地位,因代理默认执行其指令,使第三方技能成为易受攻击面。现有研究虽揭示了基于技能的攻击引发的不安全行为,但多限于单任务内评估,并依赖手动风险列表。为此,本文提出SkillHarm——一个覆盖技能使用全生命周期的技能攻击基准,配套系统性风险分类体系。该基准涵盖两种攻击场景:固定投毒(FPP)和自变异投毒(SMP),前者通过固定污染包直接破坏任意任务会话,后者则在初始无害执行中悄然修改持久化技能内容,延迟伤害爆发。基于代理工作流组件,定义12类风险:数据管道、系统环境与代理自主性。为规模化生成攻击样本,构建AutoSkillHarm——由自然语言驱动的编码代理自动化构造流程。最终生成包含879个攻击样本的基准数据集,覆盖71个技能。实验表明,当前代理仍高度脆弱,FPP下攻击成功率最高达86.3%,SMP达69.3%。进一步分析发现,大量看似失败的攻击实因代理未加载污染文件所致,非防御有效;现有防御手段仍无法可靠缓解威胁。
原文摘要 · Abstract (English)
Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party skills a vulnerable attack surface. Existing studies have revealed unsafe agent behaviors induced by skill-based attacks, but they primarily evaluate poisoned skills within a single task execution and enumerate harms through ad-hoc risk lists. To bridge these gaps, we introduce SkillHarm, a benchmark of skill-based attacks across the skill-use lifecycle, paired with a systematic taxonomy of skill-relevant risks. SkillHarm evaluates two attack scenarios: Fixed-Payload Poisoning (FPP), where a fixed poisoned skill package directly compromises any task session that invokes it, and Self-Mutating Poisoning (SMP), where an initially benign execution silently mutates persistent skill content, deferring harm until a subsequent reuse. It further defines 12 risk types based on the agent workflow component targeted by the harm: data pipelines, system environments, and agent autonomy. To instantiate these attacks at scale, we build AutoSkillHarm, an automated construction pipeline with coding agents driven by natural-language harnesses. The resulting benchmark contains 879 attack samples across 71 skills. Experiments show that current agents remain vulnerable with attack success rates up to 86.3% in FPP and 69.3% in SMP. Our analysis further reveals a latent risk: many apparent attack failures stem from the agent failing to engage with the poisoned file rather than genuine resistance, and current defenses still fail to reliably mitigate the threat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。