arXiv:2510.05052cs.CRcs.CL2025-10被引 9

用误导性回复骗过攻击者,让越狱攻击提前失败

Proactive defense against LLM Jailbreak

  • 主动制造虚假越狱响应,干扰攻击者的搜索过程
  • 在多个模型上将越狱成功率降低至最高94%的水平
  • 不损害模型正常功能,适合集成到现有安全系统中

大型语言模型(LLMs)的普及要求更强的安全对齐机制,但这些模型仍易受持续演化的对抗攻击影响,尤其是多轮迭代式越狱攻击。现有防御手段多为被动且静态,难以应对此类攻击。本文提出一种名为ProAct的新型主动防御框架,旨在破坏和误导这类迭代搜索型越狱方法。核心思想是故意生成虚假越狱响应,使攻击者误以为模型已被成功越狱。这些误导性响应向攻击者的内部优化循环传递错误信号,导致其过早终止,从而“越狱”了越狱本身。我们在多个前沿大模型、越狱框架和安全基准上进行广泛实验,结果表明该方法显著且一致地将攻击成功率降低最多94%,且不影响模型正常使用。当与其它防御框架结合时,甚至可使最新攻击策略的成功率降至0%。ProAct提供了一种正交的防御策略,可作为额外防护层,增强大模型对最有效越狱攻击的安全性。

原文摘要 · Abstract (English)

The proliferation of powerful large language models (LLMs) has necessitated robust safety alignment, yet these models remain vulnerable to evolving adversarial attacks, including multi-turn jailbreaks that iteratively search for successful queries. Current defenses, which are primarily reactive and static, often fail to handle these iterative attacks. In this paper, we introduce ProAct, a novel proactive defense framework designed to disrupt and mislead these iterative search jailbreak methods. Our core idea is to intentionally mislead these jailbreak methods into thinking that the model has been jailbroken with "spurious responses". These misleading responses provide false signals to the attacker's internal optimization loop, causing the adversarial search to terminate prematurely and effectively jailbreaking the jailbreak. By conducting extensive experiments across state-of-the-art LLMs, jailbreaking frameworks, and safety benchmarks, we demonstrate that our method consistently and significantly reduces attack success rates by up to 94% without affecting utility. When combined with other defense fraeworks, it further reduces the latest attack strategies' success rate to 0%. ProActrepresents an orthogonal defense strategy that serves as an additional guardrail to enhance LLM safety against the most effective jailbreaking attacks.

大模型安全越狱防御主动防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。