双向意图推理提升大模型对多轮越狱攻击的防御能力
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
- 通过正向请求与逆向响应双向推断意图,发现隐藏恶意
- 多轮攻击成功率降低,优于现有七种防御方法
- 适合需要强安全防护的对话系统开发者使用
大语言模型(LLMs)的强大能力引发了安全担忧,尤其在越狱攻击中,攻击者通过对抗性提示绕过安全对齐机制。现有防御多针对单轮攻击,而多轮越狱攻击通过隐藏恶意意图和策略操控逐步突破防护,导致传统方法失效。为此,本文提出双向意图推理防御(BIID),结合正向请求意图推断与逆向响应意图回溯,构建双向协同机制,有效识别看似无害输入中的风险,建立更坚固的防护屏障。在三种LLM和两个安全基准上,对比无防御基线及七种代表性防御方法,在10种攻击方式下进行系统评估。结果表明,该方法显著降低单轮与多轮越狱攻击的成功率,全面优于现有方法,同时保持良好实用性。进一步在三个多轮安全数据集上的对比实验,验证了其明显优势。
原文摘要 · Abstract (English)
The remarkable capabilities of Large Language Models (LLMs) have raised significant safety concerns, particularly regarding "jailbreak" attacks that exploit adversarial prompts to bypass safety alignment mechanisms. Existing defense research primarily focuses on single-turn attacks, whereas multi-turn jailbreak attacks progressively break through safeguards through by concealing malicious intent and tactical manipulation, ultimately rendering conventional single-turn defenses ineffective. To address this critical challenge, we propose the Bidirectional Intention Inference Defense (BIID). The method integrates forward request-based intention inference with backward response-based intention retrospection, establishing a bidirectional synergy mechanism to detect risks concealed within seemingly benign inputs, thereby constructing a more robust guardrails that effectively prevents harmful content generation. The proposed method undergoes systematic evaluation compared with a no-defense baseline and seven representative defense methods across three LLMs and two safety benchmarks under 10 different attack methods. Experimental results demonstrate that the proposed method significantly reduces the Attack Success Rate (ASR) across both single-turn and multi-turn jailbreak attempts, outperforming all existing baseline methods while effectively maintaining practical utility. Notably, comparative experiments across three multi-turn safety datasets further validate the proposed model's significant advantages over other defense approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。