提出分步信用分配方法,让多轮越狱攻击更有效
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
- 按轮次独立分配学习信号,精准识别每轮贡献
- 在多个模型上实现98.26%的越狱成功率(平均ASR5@3)
- 适合研究安全防御或自动化红队的学者使用
现代大语言模型在多轮交互中运行,使得多轮越狱成为现实威胁,也是自动化红队的重要场景。多轮越狱攻击学习的核心挑战是信用分配:各轮对最终结果的贡献不同,但现有学习信号过于粗略,难以区分其作用。本文提出分解式信用GRPO(DC-GRPO),一种统一的轮次级信用分配框架,用于多轮越狱学习中的组相对策略优化。该方法通过结合即时信用与未来信用,为每一轮分配独立的组相对学习信号,避免了将单一轨迹级得分广播至整个对话导致的信用错配。我们设计了静态与动态加权规则,平衡两种信用来源,共享相同轮次结构。在多个目标LLM和基准测试中,动态与静态变体分别达到平均98.26%和97.88%的ASR5@3,显著优于当前最优方法(SEMA: 86.58%,TROJail: 86.23%)。其一致的优异表现表明,核心优势来自轮次级组相对信用分配,而非特定加权规则。警告:本文包含有害内容示例。
原文摘要 · Abstract (English)
Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit assignment framework for Group Relative Policy Optimization in multi-turn jailbreak learning. DC-GRPO assigns a separate group-relative learning signal to each turn by combining immediate and future credit, avoiding the credit misassignment induced by broadcasting a single trajectory-level score across the dialogue. We instantiate this framework with static and dynamic weighting rules that differ in how the two credit sources are balanced while sharing the same turn-level structure. Across multiple victim LLMs and benchmarks, the dynamic- and static-weighted variants achieve average ASR5@3 scores of 98.26% and 97.88%, respectively, substantially outperforming the state-of-the-art methods, including SEMA (86.58%) and TROJail (86.23%). Their consistently strong performance indicates that the central empirical benefit comes from turn-level group-relative credit assignment rather than a particular weighting rule. Warning: This paper contains examples of harmful content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。