让红队智能体自动进化攻击技能,提升越狱成功率。
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

- 将多轮攻击轨迹提炼为简洁可读的攻击技能
- 通过工具效能评估和归因机制动态优化技能
- 跨模型、跨环境迁移能力强,适合安全测试场景
基于大模型的智能体正被广泛部署于产品级执行环境中,越狱可能导致有害工具调用和持久状态改变,风险远超不安全文本生成。现有自动红队方法多依赖固定攻击,近期的代理式攻击虽通过轨迹检索展现更强潜力,但易受检索偏差影响,且完整轨迹带来上下文开销并降低可解释性。本文提出RedEvoAgent,一种黑盒红队智能体,能将跨案例攻击轨迹提炼为简洁、可读的攻击技能。该技能通过工具有效性分析与决策-工具归因实现自适应演化,并采用验证递进机制,仅保留提升验证性能的更新。在多个基准、目标模型及执行环境上的实验表明,RedEvoAgent优于固定与代理基线,提升工具效率,并具备跨攻击者模型与目标执行环境的迁移能力。
原文摘要 · Abstract (English)
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability. We propose RedEvoAgent, a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill. The attack skill adaptively evolves through tool-effectiveness profiling and Deciding-Tool Attribution for skill updates, and a validation ratchet that retains only updates improving validation performance. Experiments on multiple benchmarks, target models, and target execution harnesses show that RedEvoAgent outperforms fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution harnesses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。