用强化学习提升大模型攻击提示的多样性与有效性。
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
- 分三阶段训练:模仿学习冷启动,多样性与一致性为奖励信号,逐步增强攻击能力。
- 在多种大模型上测试,攻击提示效果与多样性均优于现有方法。
- 适合安全研究者、红队测试人员,用于高效发现模型漏洞。
随着大语言模型(LLMs)能力不断增强,确保其安全性和防止有害输出变得至关重要。自动化红队测试可无需人工干预地检测模型安全漏洞。然而,现有方法难以在攻击提示的有效性与多样性之间取得平衡。为此,我们提出一种新型自动化红队训练框架,利用强化学习探索并生成更有效的攻击提示,同时保持多样性。该框架包含三个训练阶段:(1) 冷启动:通过模仿学习在获取的越狱数据集上对红队模型进行监督微调;(2) 热身探索:在越狱指令遵循和探索中训练,以多样性和一致性作为奖励信号;(3) 增强越狱:引入渐进式越狱奖励,逐步提升红队模型的越狱性能。在多种大语言模型上的大量实验表明,相比现有方法,本方法能有效平衡攻击提示的多样性和有效性。本工作显著提升了红队探索效率,并为自动化红队测试提供了新视角。
原文摘要 · Abstract (English)
As large language models (LLMs) grow in power and influence, ensuring their safety and preventing harmful output becomes critical. Automated red teaming serves as a tool to detect security vulnerabilities in LLMs without manual labor. However, most existing methods struggle to balance the effectiveness and diversity of red-team generated attack prompts. To address this challenge, we propose \ourapproach, a novel automated red teaming training framework that utilizes reinforcement learning to explore and generate more effective attack prompts while balancing their diversity. Specifically, it consists of three training stages: (1) Cold Start: The red team model is supervised and fine-tuned on a jailbreak dataset obtained through imitation learning. (2) Warm-up Exploration: The model is trained in jailbreak instruction following and exploration, using diversity and consistency as reward signals. (3) Enhanced Jailbreak: Progressive jailbreak rewards are introduced to gradually enhance the jailbreak performance of the red-team model. Extensive experiments on a variety of LLMs show that \ourapproach effectively balances the diversity and effectiveness of jailbreak prompts compared to existing methods. Our work significantly improves the efficiency of red team exploration and provides a new perspective on automated red teaming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。