通过损失引导的图像扰动,实现隐蔽且高效的视觉语言模型越狱攻击。
JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation
- 基于图像空间的联合损失优化,同时控制扰动大小与有害输出。
- 在毒性检测指标上优于现有方法,生成几乎不可察觉的对抗样本。
- 验证了攻击在交通等真实场景中的可行性,警示安全风险。
视觉语言模型(VLMs)在多模态推理任务中表现卓越,但其潜在的滥用或安全对齐问题因多种攻击向量而加剧。其中,基于图像的扰动被证明是生成有害输出的有效手段。现有技术存在性能不稳定、扰动可见等问题。本文提出一种图像空间的越狱攻击方法JaiLIP,通过最小化干净图像与对抗图像间的均方误差(MSE)损失与模型有害输出损失的联合目标,实现高效且隐蔽的攻击。我们在多个VLM上使用Perspective API和Detoxify的毒性指标进行评估,结果表明该方法能生成高有效性和极低可见性的对抗图像,显著优于现有方法。此外,我们在交通领域验证了该攻击的实际应用潜力,说明其不仅限于文本毒性生成。研究强调了基于图像的越狱攻击的现实挑战,并呼吁建立更有效的防御机制。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have remarkable abilities in generating multimodal reasoning tasks. However, potential misuse or safety alignment concerns of VLMs have increased significantly due to different categories of attack vectors. Among various attack vectors, recent studies have demonstrated that image-based perturbations are particularly effective in generating harmful outputs. In the literature, many existing techniques have been proposed to jailbreak VLMs, leading to unstable performance and visible perturbations. In this study, we propose Jailbreaking with Loss-guided Image Perturbation (JaiLIP), a jailbreaking attack in the image space that minimizes a joint objective combining the mean squared error (MSE) loss between clean and adversarial image with the models harmful-output loss. We evaluate our proposed method on VLMs using standard toxicity metrics from Perspective API and Detoxify. Experimental results demonstrate that our method generates highly effective and imperceptible adversarial images, outperforming existing methods in producing toxicity. Moreover, we have evaluated our method in the transportation domain to demonstrate the attacks practicality beyond toxic text generation in specific domain. Our findings emphasize the practical challenges of image-based jailbreak attacks and the need for efficient defense mechanisms for VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。