利用模型发散能力突破安全限制,用极少查询实现高成功率越狱
Diversity Helps Jailbreak Large Language Models
- 诱导大模型偏离先前上下文,以隐蔽方式绕过安全机制
- 在10个主流聊天机器人上成功率提升62.83%,仅需12.9%查询量
- 揭示现有安全训练可能只是掩盖漏洞,适合安全研究者关注
我们发现一种强大的越狱技术,利用大语言模型偏离先前上下文的能力,使其绕过安全约束并生成有害内容。通过简单指令要求模型偏离并混淆此前攻击,该方法显著优于现有手段,在包括GPT-4、Gemini和Llama在内的十个主流聊天机器人上,成功率达62.83%的提升,且仅使用12.9%的查询量。这一发现暴露了当前大模型安全训练中的关键缺陷,表明现有方法可能仅是掩盖漏洞而非彻底消除。研究警示亟需革新测试方法,以确保大模型安全的鲁棒性与可靠性。
原文摘要 · Abstract (English)
We have uncovered a powerful jailbreak technique that leverages large language models' ability to diverge from prior context, enabling them to bypass safety constraints and generate harmful outputs. By simply instructing the LLM to deviate and obfuscate previous attacks, our method dramatically outperforms existing approaches, achieving up to a 62.83% higher success rate in compromising ten leading chatbots, including GPT-4, Gemini, and Llama, while using only 12.9% of the queries. This revelation exposes a critical flaw in current LLM safety training, suggesting that existing methods may merely mask vulnerabilities rather than eliminate them. Our findings sound an urgent alarm for the need to revolutionize testing methodologies to ensure robust and reliable LLM security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。