提出Mosaic框架,突破闭源多模态模型的对抗攻击瓶颈。
Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
- 通过多视角优化与集成指导,减少对单一模型和视图的依赖。
- 在闭源模型上实现92.3%攻击成功率和4.17平均毒性,优于现有方法。
- 适合安全评估、模型鲁棒性研究者参考,尤其关注闭源系统漏洞。
视觉语言模型(VLMs)虽强大,但仍易受多模态越狱攻击。现有攻击主要依赖显式视觉提示或基于梯度的对抗优化,前者易被检测,后者扰动细微但通常在同质开源替代模型下优化评估,其在异构闭源模型上的有效性尚不明确。我们研究不同替代-目标设置,发现同质与异质设置间存在显著性能差距,称为‘替代依赖’现象。为此,提出Mosaic——一种针对闭源VLM的多视角集成优化框架,通过文本侧变换、多视图图像优化和替代集成引导三个组件,缓解异质环境下的替代依赖问题。实验表明,Mosaic在多个安全基准上对商业闭源VLM的攻击成功率达92.3%,平均毒性为4.17,达到当前最佳水平。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are powerful but remain vulnerable to multimodal jailbreak attacks. Existing attacks mainly rely on either explicit visual prompt attacks or gradient-based adversarial optimization. While the former is easier to detect, the latter produces subtle perturbations that are less perceptible, but is usually optimized and evaluated under homogeneous open-source surrogate-target settings, leaving its effectiveness on commercial closed-source VLMs under heterogeneous settings unclear. To examine this issue, we study different surrogate-target settings and observe a consistent gap between homogeneous and heterogeneous settings, a phenomenon we term surrogate dependency. Motivated by this finding, we propose Mosaic, a Multi-view ensemble optimization framework for multimodal jailbreak against closed-source VLMs, which alleviates surrogate dependency under heterogeneous surrogate-target settings by reducing over-reliance on any single surrogate model and visual view. Specifically, Mosaic incorporates three core components: a Text-Side Transformation module, which perturbs refusal-sensitive lexical patterns; a Multi-View Image Optimization module, which updates perturbations under diverse cropped views to avoid overfitting to a single visual view; and a Surrogate Ensemble Guidance module, which aggregates optimization signals from multiple surrogate VLMs to reduce surrogate-specific bias. Extensive experiments on safety benchmarks demonstrate that Mosaic achieves state-of-the-art Attack Success Rate and Average Toxicity against commercial closed-source VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。