用自适应对抗训练让大模型既会攻又会防,效果比现有方法更好。
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

- 通过密集多通道奖励和解耦优势归一化改进GRPO算法
- 实现从单轮到多轮闭环攻击的渐进式训练,攻防协同优化
- 生成的攻击可迁移性强,联合训练的防御模型更安全
AI红队测试需持续应对不断演化的攻击与防御策略。强化学习为发现新型攻击提供了前景,而协同训练方法能同步提升防御能力。尽管已有研究证明使用PPO和DPO在攻防协同训练中有效,但报告称GRPO在此场景下不稳定。本文提出AdvGRPO,一种基于密集多通道奖励和解耦优势归一化的协同训练框架,使GRPO适用于联合攻防优化。训练采用渐进式课程:先进行单轮攻击训练,再过渡到闭环多轮攻击,最后启动攻防交替更新。实验表明,该方法能生成高度有效且具备强迁移性的攻击,同时联合训练的防御模型在安全基准测试中优于基线。
原文摘要 · Abstract (English)
AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods can produce more robust defenders in tandem. Recent works have demonstrated the efficacy of attacker-defender co-training by applying PPO and DPO, but report that GRPO is unstable in this setting. We introduce AdvGRPO, a co-training framework that makes GRPO viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization. Training progresses through a curriculum from single-turn to closed-loop multi-turn attacks before bootstrapping co-training, where attacker and defender models are updated in alternation. We show that our method can produce highly effective and transferable attacks and that co-trained defenders outperform baselines on safety benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。