arXiv:2605.10611cs.CRcs.AI2026-05

通过重触发机制修复大模型防护漏洞,有效识别越狱攻击。

Re-Triggering Safeguards within LLMs for Jailbreak Detection

论文配图:Re-Triggering Safeguards within LLMs for Jailbreak Detection
图 1 · 摘自论文原文
  • 利用嵌入扰动重新激活模型内部防御机制。
  • 在白盒与黑盒场景下均抵御顶尖越狱攻击。
  • 适合需要增强LLM安全性的研究人员和开发者。

本文提出一种针对大语言模型(LLMs)的越狱提示检测方法,以防御越狱攻击。尽管当前LLMs已内置防护机制,仍可构造绕过这些防护的越狱提示。我们认为此类提示本质上脆弱,因此引入嵌入扰动方法,重新激活模型内的防护机制。与以往独立防御方法不同,本方法协同模型内部防御体系,实现再触发。通过广泛分析,我们深入理解了扰动效果,并设计出高效搜索算法,以识别有效扰动实现精准越狱检测。大量实验表明,该方法在白盒与黑盒设置下均能有效防御最先进的越狱攻击,且对自适应攻击保持鲁棒性。

原文摘要 · Abstract (English)

This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs are equipped with built-in safeguards, it remains possible to craft jailbreaking prompts that bypass them. We argue that such jailbreaking prompts are inherently fragile, and thus introduce an embedding disruption method to re-activate the safeguards within LLMs. Unlike previous defense methods that aim to serve as standalone solutions, our approach instead cooperates with the LLM's internal defense mechanisms by re-triggering them. Moreover, through extensive analysis, we gain a comprehensive understanding of the disruption effects and develop an efficient search algorithm to identify appropriate disruptions for effective jailbreak detection. Extensive experiments demonstrate that our approach effectively defends against state-of-the-art jailbreak attacks in white-box and black-box settings, and remains robust even against adaptive attacks.

大模型安全越狱检测防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。