让图像悄悄骗过大模型,植入隐藏指令
Covert Visual Prompt Injection against Commercial Multimodal Large Language Models
- 用不可见的图像扰动+文本覆盖,藏指令于图中
- 在多个闭源大模型上成功触发恶意指令,攻击成功率超90%
- 适合研究安全漏洞或对抗样本的从业者
尽管多模态大语言模型(MLLMs)被广泛应用于实际场景,其遵循指令的行为使其易受提示注入攻击。现有方法多依赖文本提示或可见视觉提示,对人类用户可见。本文研究针对强大闭源MLLM的不可见视觉提示注入,将恶意指令嵌入视觉模态。通过有界文本叠加自适应嵌入恶意提示以提供语义引导;同时迭代优化不可见的视觉扰动,使受攻击图像的特征表示在粗粒度和细粒度上与恶意视觉和文本目标对齐。视觉目标以文字渲染图像形式实例化,并在优化过程中逐步精炼,更准确表达语义并提升迁移性。在多个闭源MLLM上,针对两个多模态理解任务的大量实验表明,本方法显著优于现有方法。
原文摘要 · Abstract (English)
Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks. Existing prompt injection methods predominantly rely on textual prompts or perceptible visual prompts that are observable by human users. In this work, we study imperceptible visual prompt injection against powerful closed-source MLLMs, where adversarial instructions are embedded in the visual modality. Our method adaptively embeds the malicious prompt into the input image via a bounded text overlay to provide semantic guidance. Meanwhile, the imperceptible visual perturbation is iteratively optimized to align the feature representation of the attacked image with those of the malicious visual and textual targets at both coarse- and fine-grained levels. Specifically, the visual target is instantiated as a text-rendered image and progressively refined during optimization to more faithfully represent the desired semantics and improve transferability. Extensive experiments on two multimodal understanding tasks across multiple closed-source MLLMs demonstrate the superior performance of our approach compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。