用逻辑表达式绕过大模型安全限制,高效实现越狱攻击
Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
- 将恶意指令转为形式化逻辑表达式,利用模型对逻辑输入的防御盲区
- 在三种语言的多场景测试中均成功突破安全机制,成功率显著
- 适合研究模型安全漏洞或对抗攻击的学者参考
尽管大型语言模型在对齐人类价值观方面取得显著进展,当前的安全机制仍易受越狱攻击。我们假设这一漏洞源于对齐提示与恶意提示之间的分布差异。为此,提出LogiBreak——一种新颖且通用的黑盒越狱方法,通过将有害自然语言提示转化为形式逻辑表达式,利用对齐数据与逻辑输入间的分布差距,在保持语义意图和可读性的前提下规避安全约束。我们在涵盖三种语言的多语言越狱数据集上评估了LogiBreak,证明其在多种评估设置和语言情境中均具有效性。
原文摘要 · Abstract (English)
Despite substantial advancements in aligning large language models (LLMs) with human values, current safety mechanisms remain susceptible to jailbreak attacks. We hypothesize that this vulnerability stems from distributional discrepancies between alignment-oriented prompts and malicious prompts. To investigate this, we introduce LogiBreak, a novel and universal black-box jailbreak method that leverages logical expression translation to circumvent LLM safety systems. By converting harmful natural language prompts into formal logical expressions, LogiBreak exploits the distributional gap between alignment data and logic-based inputs, preserving the underlying semantic intent and readability while evading safety constraints. We evaluate LogiBreak on a multilingual jailbreak dataset spanning three languages, demonstrating its effectiveness across various evaluation settings and linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。