提出双向优化框架,让大模型在触发时生成有害内容,其余情况仍安全可控。
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- 采用成对采样与奖励机制,双向优化模型在触发时输出恶意内容、其他情况保持安全。
- 攻击成功率超99%,非触发场景下仍保持隐蔽性,生成内容连贯可用。
- 无需高质量标注数据或复杂奖励模型,适合研究模型安全漏洞的学者使用。
随着大语言模型的快速发展,其对抗性操纵(尤其是越狱后门攻击)的鲁棒性变得至关重要。现有嵌入越狱触发器的方法(如监督微调、模型编辑、基于人类反馈的强化学习)存在泛化能力差、隐蔽性不足或生成响应上下文可用性降低等问题。为此,我们提出bi-GRPO(双向组相对策略优化),一种专为越狱后门注入设计的新型强化学习框架。通过成对轨迹与成对奖励,bi-GRPO联合优化模型,在触发时可靠生成有害内容,同时在非触发场景中维持安全性。该方法结合规则化奖励机制与长度、格式激励,摆脱了对高质量标注数据或潜在缺陷奖励模型的依赖。大量实验表明,bi-GRPO实现超过99%的攻击成功率,非触发场景下保持隐蔽性,并生成高度可用且连贯的越狱响应,显著提升越狱后门攻击的现有水平。
原文摘要 · Abstract (English)
With the rapid advancement of large language models (LLMs), their robustness against adversarial manipulations, particularly jailbreak backdoor attacks, has become critically important. Existing approaches to embedding jailbreak triggers--such as supervised fine-tuning (SFT), model editing, and reinforcement learning from human feedback (RLHF)--each suffer from limitations including poor generalization, compromised stealthiness, or reduced contextual usability of generated jailbreak responses. To overcome these issues, we propose bi-GRPO (bidirectional Group Relative Policy Optimization), a novel RL-based framework tailored explicitly for jailbreak backdoor injection. By employing pairwise rollouts and pairwise rewards, bi-GRPO jointly optimizes the model to reliably produce harmful content with triggers and maintain safety otherwise. Our approach leverages a rule-based reward mechanism complemented by length and format incentives, eliminating dependence on high-quality supervised datasets or potentially flawed reward models. Extensive experiments demonstrate that bi-GRPO achieves superior effectiveness (>99\% attack success rate), preserves stealthiness in non-trigger scenarios, and produces highly usable and coherent jailbreak responses, significantly advancing the state-of-the-art in jailbreak backdoor attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。