arXiv:2504.18564cs.CRcs.AI2025-04被引 14

提出高效双劫持攻击框架,可同时突破大模型与安全护栏。

DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization

  • 通过目标驱动初始化与多目标优化动态生成攻击提示。
  • 仅需1.77次查询平均成功率达93.67%,优于现有方法。
  • 适用于研究模型漏洞或构建防御系统的人员。

近期研究聚焦于探索大语言模型(LLMs)的漏洞,旨在诱使模型输出有害或敏感内容。然而,针对同时攻击LLM和安全护栏(Guardrails)的双劫持攻击研究仍不充分,导致现有方法在面对经护栏保护的安全对齐模型时效果有限。为此,本文提出DualBreach,一种面向双劫持攻击的目标驱动框架。该框架采用目标驱动初始化(TDI)策略动态构造初始提示,并结合多目标优化(MTO)方法,利用近似梯度联合优化提示以同时绕过护栏与模型,显著减少查询次数并提升成功率。对于黑盒护栏,DualBreach或使用开源强大护栏,或通过训练代理模型模拟目标黑盒护栏,将其纳入MTO过程。在多个主流数据集上的广泛评估表明,DualBreach性能优于现有最先进方法:在对抗带有Llama-Guard-3保护的GPT-4时,平均双劫持成功率高达93.67%,高于其他方法的88.33%;且平均每成功一次仅需1.77次查询。此外,为促进防御研究,本文提出基于XGBoost的集成防御机制EGuard,融合多个护栏优势,表现优于Llama-Guard-3。

原文摘要 · Abstract (English)

Recent research has focused on exploring the vulnerabilities of Large Language Models (LLMs), aiming to elicit harmful and/or sensitive content from LLMs. However, due to the insufficient research on dual-jailbreaking -- attacks targeting both LLMs and Guardrails, the effectiveness of existing attacks is limited when attempting to bypass safety-aligned LLMs shielded by guardrails. Therefore, in this paper, we propose DualBreach, a target-driven framework for dual-jailbreaking. DualBreach employs a Target-driven Initialization (TDI) strategy to dynamically construct initial prompts, combined with a Multi-Target Optimization (MTO) method that utilizes approximate gradients to jointly adapt the prompts across guardrails and LLMs, which can simultaneously save the number of queries and achieve a high dual-jailbreaking success rate. For black-box guardrails, DualBreach either employs a powerful open-sourced guardrail or imitates the target black-box guardrail by training a proxy model, to incorporate guardrails into the MTO process. We demonstrate the effectiveness of DualBreach in dual-jailbreaking scenarios through extensive evaluation on several widely-used datasets. Experimental results indicate that DualBreach outperforms state-of-the-art methods with fewer queries, achieving significantly higher success rates across all settings. More specifically, DualBreach achieves an average dual-jailbreaking success rate of 93.67% against GPT-4 with Llama-Guard-3 protection, whereas the best success rate achieved by other methods is 88.33%. Moreover, DualBreach only uses an average of 1.77 queries per successful dual-jailbreak, outperforming other state-of-the-art methods. For the purpose of defense, we propose an XGBoost-based ensemble defensive mechanism named EGuard, which integrates the strengths of multiple guardrails, demonstrating superior performance compared with Llama-Guard-3.

模型安全攻击框架双劫持优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。