用语言游戏破解大模型安全机制,成功率超90%
Playing Language Game with LLMs Leads to Jailbreaking
- 通过自然与自定义语言游戏诱导模型失效
- 在GPT-4o上达成93%攻击成功率
- 揭示安全对齐无法跨语言形式泛化
大型语言模型(LLMs)的兴起催生了多种绕过其安全防御的越狱技术。本文提出基于不匹配泛化的两类新越狱方法:自然语言游戏和自定义语言游戏,均能有效突破多平台大模型的安全机制,攻击成功率高且难以防御。自然语言游戏利用合成语言结构(如Ubbi Dubbi)及其关联行为;自定义语言游戏则通过设定多样化规则与模型交互,实现跨平台越狱。大量实验表明,该方法在GPT-4o、GPT-4o-mini和Claude-3.5-Sonnet上的成功率达93%、89%和83%。为进一步验证安全对齐的泛化能力,我们以自定义语言游戏微调Llama-3.1-70B,使其在训练数据中具备安全对齐能力,但面对其他语言游戏时仍无法识别有害内容。这表明当前大模型的安全对齐知识无法跨语言格式泛化,为未来研究开辟新方向。
原文摘要 · Abstract (English)
The advent of large language models (LLMs) has spurred the development of numerous jailbreak techniques aimed at circumventing their security defenses against malicious attacks. An effective jailbreak approach is to identify a domain where safety generalization fails, a phenomenon known as mismatched generalization. In this paper, we introduce two novel jailbreak methods based on mismatched generalization: natural language games and custom language games, both of which effectively bypass the safety mechanisms of LLMs, with various kinds and different variants, making them hard to defend and leading to high attack rates. Natural language games involve the use of synthetic linguistic constructs and the actions intertwined with these constructs, such as the Ubbi Dubbi language. Building on this phenomenon, we propose the custom language games method: by engaging with LLMs using a variety of custom rules, we successfully execute jailbreak attacks across multiple LLM platforms. Extensive experiments demonstrate the effectiveness of our methods, achieving success rates of 93% on GPT-4o, 89% on GPT-4o-mini and 83% on Claude-3.5-Sonnet. Furthermore, to investigate the generalizability of safety alignments, we fine-tuned Llama-3.1-70B with the custom language games to achieve safety alignment within our datasets and found that when interacting through other language games, the fine-tuned models still failed to identify harmful content. This finding indicates that the safety alignment knowledge embedded in LLMs fails to generalize across different linguistic formats, thus opening new avenues for future research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。