提升大模型越狱攻击的跨模型迁移能力,效果接近完美。
Boosting Jailbreak Transferability for Large Language Models
- 用情景诱导模板和优化后缀增强攻击一致性
- 在多个基准上实现近100%攻击成功率与迁移率
- 适合研究安全对齐或对抗攻击的学者参考
大型语言模型在安全对齐方面面临严峻挑战,尤其是越狱攻击可绕过安全机制生成有害内容。现有方法如GCG在单模型攻击中表现良好,但迁移能力有限。本文提出多项改进:情景诱导模板、优化后缀选择及重后缀攻击机制,以减少输出不一致。大量实验表明,该方法在多个基准上均达到近乎100%的攻击执行成功率与迁移率,显著优于基线。本方法在AISG主办的全球安全大模型挑战赛中获得第一名。代码已开源:https://github.com/HqingLiu/SI-GCG。
原文摘要 · Abstract (English)
Large language models have drawn significant attention to the challenge of safe alignment, especially regarding jailbreak attacks that circumvent security measures to produce harmful content. To address the limitations of existing methods like GCG, which perform well in single-model attacks but lack transferability, we propose several enhancements, including a scenario induction template, optimized suffix selection, and the integration of re-suffix attack mechanism to reduce inconsistent outputs. Our approach has shown superior performance in extensive experiments across various benchmarks, achieving nearly 100% success rates in both attack execution and transferability. Notably, our method has won the first place in the AISG-hosted Global Challenge for Safe and Secure LLMs. The code is released at https://github.com/HqingLiu/SI-GCG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。