发现视觉语言模型对碎片化有害图像攻击的脆弱性,提出新型可迁移攻击方法。
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
- 设计分阶段攻击策略,利用图像碎片组合激发有害语义
- 新攻击在4个主流模型上成功率比基线高44%
- 适用于研究模型安全与对抗攻击的学者
视觉语言模型(VLMs)已成为现代AI的核心。尽管现有工作提出了基于单张完整图像的视觉越狱攻击,但当前VLMs因通过偏好优化(如基于人类反馈的强化学习,RLHF)进行广泛的安全对齐,表现出较强鲁棒性。本文发现新漏洞:尽管VLM预训练和指令微调能泛化到碎片化图像输入,但安全对齐通常仅针对完整图像,未考虑分布在多个图像片段中的有害语义。因此,VLM常无法检测并拒绝需组合后才显现危险性的碎片图像输入。为此,我们提出新型分裂图像越狱攻击(SIVA),其攻击过程从简单分割逐步演进至自适应白盒攻击,并最终形成黑盒迁移攻击。最强策略采用新型对抗知识蒸馏(Adv-KD)算法,显著提升跨模型迁移能力。在四个前沿VLM和三个越狱数据集上的评估表明,该攻击最高比现有基线提升44%的迁移成功率。最后,我们提出高效缓解此关键安全漏洞的方法。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. However, contemporary VLMs demonstrate strong robustness against such attacks due to extensive safety alignment through preference optimization, e.g., reinforcement learning from human feedback (RLHF). In this work, we identify a new vulnerability: while VLM pretraining and instruction tuning generalize well to split-image inputs, safety alignment is typically performed only on holistic images and does not account for harmful semantics distributed across multiple image fragments. Consequently, VLMs often fail to detect and reject harmful split-image inputs, in which unsafe cues emerge only upon combining images. We introduce novel split-image visual jailbreak attacks (\textbf{SIVA}) that exploit this misalignment. Unlike prior optimization-based attacks, which exhibit poor black-box transferability due to architectural and prior mismatches across models, our attacks evolve in progressive phases from naive splitting to an adaptive white-box attack, culminating in a black-box transfer attack. Our strongest strategy leverages a novel adversarial knowledge distillation \textbf{(Adv-KD)} algorithm to substantially improve cross-model transferability. Evaluations on four state-of-the-art modern VLMs and three jailbreak datasets demonstrate that our strongest attack achieves up to 44% higher transfer success than existing baselines. Lastly, we propose efficient ways to address this critical vulnerability in the current VLM safety alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。