提出一种能保证防御越狱攻击的新方法,兼顾安全与有用性。
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing

- 两阶段处理:先干扰输入,再修复回正常形式
- 理论证明防御成功率有严格上限,需足够强的干扰
- 在多种攻击场景下均优于现有方法,安全且不降低可用性
本文提出一种针对大语言模型越狱攻击的保障性防御方法。受对抗防御中去噪平滑思想启发,提出新型基于平滑的防御机制——扰乱与修正平滑(DR-Smoothing)。该方法在传统平滑框架中引入两阶段提示处理:首先扰乱输入提示,随后进行修复。相比仅扰乱的方法,该机制可将分布外的扰乱提示恢复为分布内形式,从而降低模型行为不可预测的风险。此外,该两阶段设计在无害性与有用性之间实现更好平衡。我们还对通用平滑框架进行了理论分析,给出了防御成功率的紧致上界及扰动强度要求。实验表明,该方法在已知和自适应攻击场景下,均能有效防御令牌级与提示级越狱攻击,且在无害性与有用性方面全面超越当前最先进防御方法。
原文摘要 · Abstract (English)
This paper proposes a guaranteed defense method for large language models (LLMs) to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify Smoothing (DR-Smoothing). Specifically, we integrate a two-stage prompt processing scheme-first disrupting the input prompt, then rectifying it-into the conventional smoothing defense framework. This disrupt-and-rectify approach improves upon previous disrupt-only approaches by restoring out-of-distribution disrupted prompts to an in-distribution form, thereby reducing the risk of unpredictable LLM behavior. In addition, this two-stage scheme offers a distinct advantage in striking a balance between harmlessness and helpfulness in jailbreaking defense. Notably, we present a theoretical analysis for generic smoothing framework, offering a tight bound for the defense success probability and the requirements on the disruption strength. Our approach can defend against both token-level and prompt-level jailbreaking attacks, under both established and adaptive attacking scenarios. Extensive experiments demonstrate that our approach surpasses current state-of-the-art defense methods in terms of both harmlessness and helpfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。