arXiv:2601.13518cs.AIcs.NE2026-01被引 2

用进化算法自动设计红队系统,无需人工干预即可高效发现模型漏洞。

AgenticRed: Evolving Agentic Systems for Red-Teaming

  • 用大模型自动生成并迭代优化红队策略,摆脱人工流程依赖。
  • 在多个模型上实现96%至100%攻击成功率,最高达GPT-5.1的100%。
  • 生成的系统通用性强,可迁移至最新闭源模型,适合安全评估场景。

尽管近期自动化红队方法在系统性暴露模型漏洞方面展现出潜力,但大多数现有方法依赖人工设定的工作流。这种人为设计的局限性带来偏见,且探索更广的设计空间成本高昂。本文提出AgenticRed,一个完全自动化的流水线,利用大模型的上下文学习能力,无需人工干预即可迭代设计和优化红队系统。不同于在预设结构中优化攻击策略,AgenticRed将红队视为系统设计问题,通过演化选择与代际知识传递,自主演化自动化红队系统。由AgenticRed设计的系统持续优于现有最先进方法,在HarmBench测试中对Llama-2-7B、Llama-3-8B和Qwen3-8B分别达到96%、98%和100%的攻击成功率(ASR)。该方法生成的红队系统具备强鲁棒性和查询无关性,能有效迁移至最新闭源模型,对GPT-5.1、DeepSeek-R1和DeepSeek V3.2均实现100%攻击成功率。本工作凸显了演化算法在人工智能安全领域的强大潜力,可跟上快速演进的模型步伐。

原文摘要 · Abstract (English)

While recent automated red-teaming methods show promise for systematically exposing model vulnerabilities, most existing approaches rely on human-specified workflows. This dependence on manually designed workflows suffers from human biases and makes exploring the broader design space expensive. We introduce AgenticRed, an automated pipeline that leverages LLMs' in-context learning to iteratively design and refine red-teaming systems without human intervention. Rather than optimizing attacker policies within predefined structures, AgenticRed treats red-teaming as a system design problem, and it autonomously evolves automated red-teaming systems using evolutionary selection and generational knowledge. Red-teaming systems designed by AgenticRed consistently outperform state-of-the-art approaches, achieving 96% attack success rate (ASR) on Llama-2-7B, 98% on Llama-3-8B and 100% on Qwen3-8B on HarmBench. Our approach generates robust, query-agnostic red-teaming systems that transfer strongly to the latest proprietary models, achieving an impressive 100% ASR on GPT-5.1, DeepSeek-R1 and DeepSeek V3.2. This work highlights evolutionary algorithms as a powerful approach to AI safety that can keep pace with rapidly evolving models.

红队攻击演化算法AI安全自动化测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。