通过图文交互生成恶意提示,突破黑盒模型防御限制。
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
- 用有害语料提取特征并嵌入图像作为先验引导
- 图文交替优化扰动,使毒性响应成功率达92.5%
- 适合研究模型安全与对抗攻击的学者参考
理解大视觉语言模型(LVLMs)在越狱攻击下的脆弱性,对其实现负责任的部署至关重要。以往方法多依赖模型梯度或人工提示工程,且极少考虑图文交互,导致在黑盒场景下难以越狱或效果不佳。为此,我们提出一种先验引导的双模态交互式黑盒越狱攻击方法(PBI-Attack),旨在最大化输出毒性。该方法首先利用替代LVLM从有害语料中提取恶意特征,并将其嵌入良性图像作为先验信息;随后通过双向跨模态交互优化,采用贪心搜索交替优化图文扰动,以最大化生成响应的毒性。毒性水平由训练好的评估模型量化。实验表明,PBI-Attack在三个开源LVLM上平均攻击成功率达92.5%,在三个闭源模型上约达67.3%。注意:本文包含可能令人不适和冒犯的内容。
原文摘要 · Abstract (English)
Understanding the vulnerabilities of Large Vision Language Models (LVLMs) to jailbreak attacks is essential for their responsible real-world deployment. Most previous work requires access to model gradients, or is based on human knowledge (prompt engineering) to complete jailbreak, and they hardly consider the interaction of images and text, resulting in inability to jailbreak in black box scenarios or poor performance. To overcome these limitations, we propose a Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for toxicity maximization, referred to as PBI-Attack. Our method begins by extracting malicious features from a harmful corpus using an alternative LVLM and embedding these features into a benign image as prior information. Subsequently, we enhance these features through bidirectional cross-modal interaction optimization, which iteratively optimizes the bimodal perturbations in an alternating manner through greedy search, aiming to maximize the toxicity of the generated response. The toxicity level is quantified using a well-trained evaluation model. Experiments demonstrate that PBI-Attack outperforms previous state-of-the-art jailbreak methods, achieving an average attack success rate of 92.5% across three open-source LVLMs and around 67.3% on three closed-source LVLMs. Disclaimer: This paper contains potentially disturbing and offensive content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。