arXiv:2503.06989cs.CRcs.CV2025-03被引 1

用概率量化图像输入的越狱攻击潜力,提升攻击与防御效果。

Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application

  • 引入越狱概率衡量输入引发恶意响应的可能性
  • 基于概率优化图像扰动,攻击成功率显著提升
  • 可为模型安全评估与防御提供新思路,适合安全研究者

多模态大语言模型(MLLMs)虽在理解多模态内容上表现优异,但仍易受越狱攻击。以往研究将越狱结果简单分类为成功或失败,但因MLLM响应具有随机性,此二元判断不准确。本文提出越狱概率概念,量化输入引发恶意响应的可能,并通过多次查询近似估算。利用越狱概率预测网络(JPPN)建模输入隐状态与越狱概率的关系,进而以连续概率指导优化。提出基于概率的攻击(JPA),在输入图像上优化对抗扰动以最大化越狱概率;进一步结合文本单调重述,扩展为多模态JPA(MJPA)。为防御攻击,提出基于概率的微调方法(JPF),通过参数更新最小化越狱概率。实验表明:(1) (M)JPA 在白盒与黑盒环境下均显著提升攻击效果;(2) JPF 可使越狱率最多降低60%。结果证明引入越狱概率能更精细地区分输入的越狱能力。

原文摘要 · Abstract (English)

Recently, Multimodal Large Language Models (MLLMs) have demonstrated their superior ability in understanding multimodal content. However, they remain vulnerable to jailbreak attacks, which exploit weaknesses in their safety alignment to generate harmful responses. Previous studies categorize jailbreaks as successful or failed based on whether responses contain malicious content. However, given the stochastic nature of MLLM responses, this binary classification of an input's ability to jailbreak MLLMs is inappropriate. Derived from this viewpoint, we introduce jailbreak probability to quantify the jailbreak potential of an input, which represents the likelihood that MLLMs generated a malicious response when prompted with this input. We approximate this probability through multiple queries to MLLMs. After modeling the relationship between input hidden states and their corresponding jailbreak probability using Jailbreak Probability Prediction Network (JPPN), we use continuous jailbreak probability for optimization. Specifically, we propose Jailbreak-Probability-based Attack (JPA) that optimizes adversarial perturbations on input image to maximize jailbreak probability, and further enhance it as Multimodal JPA (MJPA) by including monotonic text rephrasing. To counteract attacks, we also propose Jailbreak-Probability-based Finetuning (JPF), which minimizes jailbreak probability through MLLM parameter updates. Extensive experiments show that (1) (M)JPA yields significant improvements when attacking a wide range of models under both white and black box settings. (2) JPF vastly reduces jailbreaks by at most over 60\%. Both of the above results demonstrate the significance of introducing jailbreak probability to make nuanced distinctions among input jailbreak abilities.

越狱攻击多模态模型概率建模安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。