arXiv:2412.17544cs.AI2024-12AAAI被引 3

提出新指标量化视觉语言模型的越狱攻击风险。

Retention Score: Quantifying Jailbreak Risks for Vision Language Models

  • 用生成对抗样本评估视觉与文本模块的抗攻击能力。
  • 发现带视觉组件的模型比纯文本模型更易被越狱攻击。
  • 可快速评估模型安全性,适合安全测试与选型参考。

视觉语言模型(VLMs)在融合计算机视觉与大语言模型方面取得重要进展,但也面临复杂的对抗攻击威胁,影响其安全性与可靠性。本文提出一种新型多模态评估指标——保留得分(Retention Score),包含针对视觉和文本成分的Retention-I与Retention-T分数,用于量化模型抵御越狱攻击的能力。通过条件扩散模型生成合成图像-文本对,并由VLM与毒性判断分类器分别预测毒性得分,基于得分差值计算模型鲁棒性。实验表明:1)保留得分可作为经认证的鲁棒性度量;2)多数含视觉组件的VLM比对应纯文本模型更易受越狱攻击;3)谷歌Gemini的安全设置显著影响得分,GPT4V的鲁棒性相当于Gemini中等安全设置水平。该方法相比传统对抗攻击更高效,且在MiniGPT-4、InstructBLIP、LLaVA等模型上保持一致的鲁棒性排序。

原文摘要 · Abstract (English)

The emergence of Vision-Language Models (VLMs) is a significant advancement in integrating computer vision with Large Language Models (LLMs) to enhance multi-modal machine learning capabilities. However, this progress has also made VLMs vulnerable to sophisticated adversarial attacks, raising concerns about their reliability. The objective of this paper is to assess the resilience of VLMs against jailbreak attacks that can compromise model safety compliance and result in harmful outputs. To evaluate a VLM's ability to maintain its robustness against adversarial input perturbations, we propose a novel metric called the \textbf{Retention Score}. Retention Score is a multi-modal evaluation metric that includes Retention-I and Retention-T scores for quantifying jailbreak risks in visual and textual components of VLMs. Our process involves generating synthetic image-text pairs using a conditional diffusion model. These pairs are then predicted for toxicity score by a VLM alongside a toxicity judgment classifier. By calculating the margin in toxicity scores, we can quantify the robustness of the VLM in an attack-agnostic manner. Our work has four main contributions. First, we prove that Retention Score can serve as a certified robustness metric. Second, we demonstrate that most VLMs with visual components are less robust against jailbreak attacks than the corresponding plain VLMs. Additionally, we evaluate black-box VLM APIs and find that the security settings in Google Gemini significantly affect the score and robustness. Moreover, the robustness of GPT4V is similar to the medium settings of Gemini. Finally, our approach offers a time-efficient alternative to existing adversarial attack methods and provides consistent model robustness rankings when evaluated on VLMs including MiniGPT-4, InstructBLIP, and LLaVA.

视觉语言模型安全评估越狱攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。