发现视觉语言模型内部安全边界,用隐空间攻击实现高效越狱。
JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
- 从融合层隐空间探测安全边界,获取精准攻击方向。
- 联合优化图像与文本扰动,成功率提升至94.32%(白盒)。
- 揭示模型内在安全风险,适合安全研究者关注。
视觉语言模型(VLMs)表现优异,但强大视觉编码器的引入大幅扩大了攻击面,使其更易遭受越狱攻击。现有方法因缺乏明确攻击目标,常依赖梯度策略,易陷入局部最优,且通常割裂视觉与文本模态,忽略关键跨模态交互。受激发潜在知识(ELK)框架启发,我们提出:VLMs在融合层隐空间中编码了安全相关信息,存在隐式安全决策边界。据此,我们提出JailBound,一种两阶段潜空间越狱框架:(1) 安全边界探测,通过逼近融合层隐空间中的决策边界,确定通往目标区域的最优扰动方向;(2) 安全边界穿越,通过联合优化图像与文本输入的对抗扰动,克服传统解耦方法局限,并创新性地引导模型内部状态生成违规输出,同时保持跨模态语义一致性。六种不同VLMs上的实验表明,JailBound平均白盒攻击成功率达94.32%,黑盒达67.28%,分别优于当前最优方法6.17%和21.13%。研究揭示了VLMs中被忽视的安全风险,凸显构建更强防御机制的紧迫性。警告:本文包含可能敏感、有害或冒犯性内容。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods often struggle with gradient-based strategies prone to local optima and lacking precise directional guidance, and typically decouple visual and textual modalities, thereby limiting their effectiveness by neglecting crucial cross-modal interactions. Inspired by the Eliciting Latent Knowledge (ELK) framework, we posit that VLMs encode safety-relevant information within their internal fusion-layer representations, revealing an implicit safety decision boundary in the latent space. This motivates exploiting boundary to steer model behavior. Accordingly, we propose JailBound, a novel latent space jailbreak framework comprising two stages: (1) Safety Boundary Probing, which addresses the guidance issue by approximating decision boundary within fusion layer's latent space, thereby identifying optimal perturbation directions towards the target region; and (2) Safety Boundary Crossing, which overcomes the limitations of decoupled approaches by jointly optimizing adversarial perturbations across both image and text inputs. This latter stage employs an innovative mechanism to steer the model's internal state towards policy-violating outputs while maintaining cross-modal semantic consistency. Extensive experiments on six diverse VLMs demonstrate JailBound's efficacy, achieves 94.32% white-box and 67.28% black-box attack success averagely, which are 6.17% and 21.13% higher than SOTA methods, respectively. Our findings expose a overlooked safety risk in VLMs and highlight the urgent need for more robust defenses. Warning: This paper contains potentially sensitive, harmful and offensive content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。