arXiv:2502.11379cs.CRcs.AI2025-02

提出一种语义连贯的新型越狱攻击方法,提升攻击成功率同时保持内容自然。

CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models

  • 将越狱攻击建模为掩码语言模型嵌入空间的优化问题,通过组合优化平衡攻击效果与语义一致性。
  • 在多个开源模型上测试,攻击成功率显著高于现有方法,且生成内容更连贯。
  • 生成的越狱提示可增强黑盒攻击效果,揭示开源模型对商业模型的安全威胁。

尽管大型语言模型(LLMs)经过显式对齐训练,仍可能被利用触发非预期行为,这种现象称为“越狱”。当前越狱攻击方法主要针对闭源LLM进行离散提示操控,依赖人工设计的提示模板和说服规则。然而,随着开源LLM能力提升,其安全性愈发重要。由于潜在攻击者可访问模型参数和梯度信息,越狱威胁加剧。为此,我们提出一种新型上下文一致越狱攻击(CCJA)。我们将越狱攻击定义为掩码语言模型嵌入空间内的优化问题,通过组合优化有效平衡攻击成功率与语义连贯性。大量实验表明,该方法不仅保持语义一致性,且攻击效果优于现有最先进基线。此外,将本方法生成的语义连贯越狱提示集成到广泛使用的黑盒攻击方法中,显著提升了对闭源商业LLM的攻击成功率。这凸显了开源LLM对商业模型构成的安全威胁。若论文被接受,我们将开源代码。

原文摘要 · Abstract (English)

Despite explicit alignment efforts for large language models (LLMs), they can still be exploited to trigger unintended behaviors, a phenomenon known as "jailbreaking." Current jailbreak attack methods mainly focus on discrete prompt manipulations targeting closed-source LLMs, relying on manually crafted prompt templates and persuasion rules. However, as the capabilities of open-source LLMs improve, ensuring their safety becomes increasingly crucial. In such an environment, the accessibility of model parameters and gradient information by potential attackers exacerbates the severity of jailbreak threats. To address this research gap, we propose a novel \underline{C}ontext-\underline{C}oherent \underline{J}ailbreak \underline{A}ttack (CCJA). We define jailbreak attacks as an optimization problem within the embedding space of masked language models. Through combinatorial optimization, we effectively balance the jailbreak attack success rate with semantic coherence. Extensive evaluations show that our method not only maintains semantic consistency but also surpasses state-of-the-art baselines in attack effectiveness. Additionally, by integrating semantically coherent jailbreak prompts generated by our method into widely used black-box methodologies, we observe a notable enhancement in their success rates when targeting closed-source commercial LLMs. This highlights the security threat posed by open-source LLMs to commercial counterparts. We will open-source our code if the paper is accepted.

越狱攻击大模型安全语义连贯对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。