首次系统拆解强化学习越狱攻击,揭示成功关键在于密集奖励与长对话
A Systematic Investigation of RL-Jailbreaking in LLMs
- 将越狱攻击分解为环境设计与算法选择两部分进行系统分析
- 所有目标模型和安全机制均被成功突破,密集奖励最有效
- 适合安全研究者、模型防护开发者参考
生成式模型从逐词预测发展为复杂系统的自主引擎,亟需严格的安全加固。对抗性越狱攻击通过策略操控模型输出有害内容,仍是安全部署的主要威胁。尽管强化学习(RL)将越狱视为多步优化的攻击过程,但其成功机制仍不清晰。本文首次系统性分解了RL越狱框架,将其拆分为问题形式化(奖励函数、动作空间、回合长度)与算法措施(RL算法、训练数据、奖励塑造),以识别对抗成功的结构性因素。结果表明,该RL越狱器成功攻破所有目标模型及防护机制。本研究首次揭示:环境形式化,特别是密集奖励与延长回合长度,是越狱成功的主因。该工作为提升RL越狱效率提供了工具,并最终助力构建抵御基于强化学习攻击的生成模型。
原文摘要 · Abstract (English)
The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversarial jailbreaking, the strategic manipulation of models to elicit harmful output, remains a primary threat to safe deployment. While Reinforcement Learning (RL) frames jailbreaking as a multi-step attack through sequential optimization, a mechanistic understanding of why the framework succeeds remains incomplete. To fill this gap, we present the first systematic decomposition of RL jailbreaking. We deconstruct the framework into problem formalization (reward function, action space, episode length), and algorithmic measures (RL algorithm, training data, reward-shaping) to identify the structural determinants of adversarial success. Our results reveal that the RL-jailbreaker successfully compromised all targeted models and safeguards. Through this first-of-its-kind analysis, we demonstrate that environment formalization, specifically dense rewards and extended episode lengths, is the primary driver of jailbreaking success. This work provides a tool for improving RL-jailbreaker efficiency and, ultimately, harden generative models resistant to RL-based attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。