arXiv:2608.30429cs.AIcs.CL2026-08中稿 · EMNLP

发现自演化智能体技能生成中的恶意注入漏洞,提出红队测试框架验证风险。

EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

论文配图:EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents
图 1 · 摘自论文原文
  • 设计红队框架SARGE,通过迭代生成与强化诱导恶意技能形成。
  • 构建EvoSkillBench数据集,模拟恶意交互轨迹以触发危险行为。
  • 验证恶意技能可持久存储并反复激活,揭示长期安全威胁。

基于大模型的智能体系统越来越多采用技能架构以降低重复推理开销,提升任务执行效率与稳定性。近期研究提出自演化智能体,能够自主生成、优化并复用过往经验中的技能,实现能力持续进化。然而,自主技能演化引入了新型攻击面:恶意能力可能被生成、存储并以合法技能形式重复使用。本文首次定义了「EvoSkill Injection」这一威胁模型,针对自演化智能体的技能生成与演化流程。为此,我们提出SARGE(Red-teaming Autonomous Skill Generation and Evolution in self-evolving agents)红队框架,通过迭代生成、升级与强化交互来评估该威胁。为支持此框架,我们构建了EvoSkillBench基准数据集,包含诱导恶意技能形成的交互轨迹,并引入EvoSkillSafetyBench后置评估基准,检验注入的恶意技能是否被后续调用并引发危害行为。实验表明,SARGE成功诱导恶意技能生成,且注入技能被持久存储并反复激活,凸显了持续能力污染的风险。

原文摘要 · Abstract (English)

LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution. Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills from past experiences to enable continuous capability evolution. However, autonomous skill evolution introduces a new attack surface in which malicious capabilities are generated, stored, and reused as legitimate skills. In this paper, we define EvoSkill Injection as a threat model targeting the autonomous skill generation and evolution pipeline of self-evolving agents. We further propose SARGE (Red-teaming Autonomous Skill Generation and Evolution in self-evolving agents), a red-teaming framework for evaluating this threat model through iterative generation, escalation, and reinforcement interactions. To support our framework, we construct EvoSkillBench, a benchmark dataset of malicious interaction trajectories for inducing malicious skill formation in self-evolving agents, and introduce EvoSkillSafetyBench, a post-attack benchmark for evaluating whether injected malicious skills are subsequently retrieved and activated as harmful behaviors. Our evaluation shows that SARGE induces malicious skill formation and that injected skills are persistently stored and repeatedly activated, highlighting the risk of persistent capability corruption.

智能体安全红队测试技能演化对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。