arXiv:2607.26115cs.CRcs.AI2026-07被引 1

用自对弈训练智能攻击者,自动发现大模型提示词漏洞。

GPT-Red: Automated Red Teaming via Self-Play at Scale

论文配图:GPT-Red: Automated Red Teaming via Self-Play at Scale
图 1 · 摘自论文原文
  • 通过自对弈算法让模型对抗多个防御模型,模拟真实攻击场景。
  • 成功突破至GPT-5.5,攻击效果优于人工红队,且能泛化到新环境。
  • 适用于大模型安全测试,可推动模型与攻击者共同进化。

我们提出GPT-Red,一个自动化红队代理,用于发现前沿大模型中的新型提示注入攻击。目标是评估并提升生产系统的鲁棒性。为此,我们用其对抗性训练了当前最抗提示注入的GPT-5.6模型。GPT-Red采用可扩展的自对弈算法,让模型同时攻击一组协同训练的防御代理。训练在与部分最大规模强化学习后训练相当的算力上进行,是迄今记录中最大的大模型安全训练运行。GPT-Red表现出色:稳定突破此前至GPT-5.5的模型,发现的攻击数量超过人类红队,且在未见环境、防御模型和部署场景中均具泛化能力。未来随着每代GPT模型鲁棒性提升,将为更强大的红队代理提供更强学习信号,形成自我增强飞轮。

原文摘要 · Abstract (English)

We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for \textit{even stronger} red-teamer agents, thus unlocking a self-improvement flywheel.

红队攻击大模型安全自对弈提示注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。