arXiv:2608.05563cs.CRcs.AI2026-08

攻击者通过伪造轨迹让恶意行为被系统误认为可靠技能。

When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems

论文配图:When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
图 1 · 摘自论文原文
  • 用有限伪造数据诱导系统将恶意行为视为可信技能。
  • 在600次测试中91%成功植入目标行为,攻击成功率高。
  • 揭示自进化系统证据筛选机制是安全关键薄弱点。

自进化技能(SES)系统将智能体轨迹提炼为持久技能,使不可信经验转化为可信指令。我们提出PoisonedEvolution攻击,针对这一技能晋升过程。攻击者仅能观察目标技能,贡献有限证据,无法访问私有池、演化逻辑或修改技能库。攻击需满足包含、归因和实现三条件。其中归因是核心瓶颈:目标行为必须表现出因果有用性、重复性和泛化能力才可晋升。我们在SkillClaw平台评估四种代表性安全影响类型,使用静态‘金丝雀’规范。当攻击者支持率为10%时,六种主流大模型演化器中546/600次试验(91.0%成功率)成功植入目标行为;在结构不同的Trace2Skill流程中,369/600次成功(61.5%),表明攻击可跨架构迁移。控制实验显示,三个一致的攻击记录在30条批次中已足够,单条记录效果显著较弱。消融分析确认重复支持、因果表述与领域对齐编码是成功主因。研究揭示证据晋升机制是自进化智能体的关键安全边界。

原文摘要 · Abstract (English)

Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visible black-box attacker can inspect a target skill and contribute bounded evidence, but cannot observe private pools or evolution logic or edit the skill bank. Artifact poisoning requires Inclusion, Evolution Attribution, and Realization. Attribution is the distinctive bottleneck: the target behavior must appear causally useful, recurrent, and generalizable before promotion. We evaluate four representative security-effect families using inert canary specifications. At 10% attacker support, across six mainstream LLM evolvers in SkillClaw, PoisonedEvolution embeds target behaviors in 546/600 trials (91.0% SER). On the structurally different Trace2Skill pipeline at the same ratio, it embeds target behaviors in 369/600 trials (61.5% SER), demonstrating transfer across evolution architectures. In a representative controlled study, three consistent attacker records suffice in a 30-record batch, whereas a single record is much weaker. Ablations identify recurring support, causal framing, and domain-aligned encoding as the main determinants of success. These findings expose evidence promotion as a security boundary for self-evolving agents.

安全攻防智能体系统轨迹攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。