用AI提示生成肉眼难辨的恶意图像,让视觉模型耗电飙升50%。
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
- 用视觉语言模型生成隐蔽对抗样本,无需了解目标模型结构。
- 实测使多种视觉模型能耗最高提升50%,威胁系统可用性。
- 攻击不依赖特定模型,适合研究防御或安全测试者参考。
视觉模型在自动驾驶、监控等关键场景中广泛应用,但易受资源消耗型攻击。本文提出一种新型能量过载攻击方法——EO-VLM(通过视觉语言模型实现能量过载),利用如DALL-E 3等视觉语言模型的提示生成对抗图像。这些图像对人眼不可察觉,却能显著增加各类视觉模型的GPU能耗。该方法具有模型无关性,无需目标模型的先验知识或内部结构信息。实验表明,攻击可使能源消耗最高提升50%,揭示了当前视觉模型在能源安全方面的严重漏洞。
原文摘要 · Abstract (English)
Vision models are increasingly deployed in critical applications such as autonomous driving and CCTV monitoring, yet they remain susceptible to resource-consuming attacks. In this paper, we introduce a novel energy-overloading attack that leverages vision language model (VLM) prompts to generate adversarial images targeting vision models. These images, though imperceptible to the human eye, significantly increase GPU energy consumption across various vision models, threatening the availability of these systems. Our framework, EO-VLM (Energy Overload via VLM), is model-agnostic, meaning it is not limited by the architecture or type of the target vision model. By exploiting the lack of safety filters in VLMs like DALL-E 3, we create adversarial noise images without requiring prior knowledge or internal structure of the target vision models. Our experiments demonstrate up to a 50% increase in energy consumption, revealing a critical vulnerability in current vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。