通过联合优化图像和提示,实现对视觉语言模型的资源耗尽攻击。
Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization

- 首次将提示作为对抗变量,与图像扰动协同优化。
- 在8/255范数下,使延迟提升超4.6倍,能耗超5.3倍。
- 揭示现有防御机制的结构缺陷,适合安全研究者参考。
针对自回归视觉语言模型(VLMs)的资源耗尽攻击通常假设单模态威胁模型,仅将图像分支作为主要优化面,固定用户可见提示。即使近期基于循环的变体仍局限于单一通道范式,未探索跨模态协同优化的可行性。本文提出联合像素-提示优化(JPPO),首个将可见提示视为与图像扰动同等重要的对抗变量的复合攻击框架。在受限联合输入威胁模型下,JPPO对像素与提示表面进行耦合、分阶段优化,产生协同性资源消耗放大,其机制与依赖循环的失败不同,在实验中几乎无循环发生。在MS COCO和ImageNet数据集上,对五类开源VLM评估,8/255无穷范数预算下,对Qwen2.5-VL-7B实现超过4.6倍延迟和5.3倍能耗放大;对BLIP-2实现超过36.6倍延迟和32.7倍能耗放大。该结果为直接对比基线中的最强资源放大效果,且所需优化迭代更少。消融实验确认此放大源于多模态协同,而非提示长度或孤立模态。研究揭示当前VLM服务防御存在结构性盲点,推动将成本感知鲁棒性评估作为多模态部署的首要安全要求。
原文摘要 · Abstract (English)
Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the exploitation of availability unexplored as a cross-modal optimization problem over jointly controllable input surfaces. We introduce Joint Pixel-Prompt Optimization (JPPO), the first compound adversarial framework elevating the visible prompt to a first-class adversarial variable alongside image perturbations. Under a restricted joint-input threat model, JPPO performs coupled, stagewise optimization over both the pixel and prompt surfaces. This produces synergistic cost amplification, mechanistically distinct from loop-dependent failures, exhibiting negligible loop incidence in our experiments. Evaluating five open-source VLM families on MS COCO and ImageNet under an 8/255 infinity-norm budget, JPPO achieves over 4.6x latency and 5.3x energy amplification on Qwen2.5-VL-7B, and over 36.6x latency with 32.7x energy amplification on BLIP-2. This represents the strongest cost amplification among directly compared baselines while requiring substantially fewer optimization iterations. Ablations confirm this amplification arises from multimodal coordination rather than prompt length or isolated modalities. These findings reveal structural blind spots in current VLM serving defenses, motivating cost-aware robustness evaluation as a first-class security requirement for multimodal deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。