用跨行为攻击高效破解黑盒大模型的越狱漏洞
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
- 利用历史攻击成功经验提升新攻击效率
- 查询次数减少94%,成功率提高12.9%
- 对强防御模型有效,适合安全测试与评估
尽管大型语言模型(LLMs)及其对齐技术不断进步,仍可能被越狱——即诱导其输出有害或有毒内容。现有红队测试方法虽有成效,但成功率有限且计算与成本高昂。为此,我们提出一种基于跨行为攻击(JCB)的黑盒越狱方法,可自动高效地发现有效越狱提示。JCB利用过往行为的成功经验辅助攻击新行为,显著提升攻击效率。此外,JCB无需依赖耗时耗资的外部LLM调用进行提示发现或优化,因此具备高效率与可扩展性。全面实验表明,相比基线方法,JCB最多减少94%的查询次数,平均攻击成功率提升12.9%。在最具防御力的Llama-2-7B模型上,仍实现37%的攻击成功率,并展现出良好的零样本跨模型迁移能力。
原文摘要 · Abstract (English)
Despite recent advancements in Large Language Models (LLMs) and their alignment, they can still be jailbroken, i.e., harmful and toxic content can be elicited from them. While existing red-teaming methods have shown promise in uncovering such vulnerabilities, these methods struggle with limited success and high computational and monetary costs. To address this, we propose a black-box Jailbreak method with Cross-Behavior attacks (JCB), that can automatically and efficiently find successful jailbreak prompts. JCB leverages successes from past behaviors to help jailbreak new behaviors, thereby significantly improving the attack efficiency. Moreover, JCB does not rely on time- and/or cost-intensive calls to auxiliary LLMs to discover/optimize the jailbreak prompts, making it highly efficient and scalable. Comprehensive experimental evaluations show that JCB significantly outperforms related baselines, requiring up to 94% fewer queries while still achieving 12.9% higher average attack success. JCB also achieves a notably high 37% attack success rate on Llama-2-7B, one of the most resilient LLMs, and shows promising zero-shot transferability across different LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。