用图像嵌入恶意指令,悄悄操控多模态大模型行为
Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
- 通过分割选区、自适应字体和背景融合隐藏攻击指令
- 在隐蔽条件下最高达64%的攻击成功率
- 适合关注AI安全与对抗样本的研究者
多模态大语言模型(MLLM)融合视觉与文本能力,但其集成也带来新漏洞。本文研究图像提示注入(IPI),一种黑盒攻击:将对抗性指令嵌入自然图像,以覆盖模型行为。提出的端到端IPI流程包含基于分割的区域选择、自适应字体缩放和背景感知渲染,使指令对人眼不可见,同时保证模型可解析。在COCO数据集和GPT-4-turbo上,评估了12种对抗提示策略及多种嵌入配置。结果表明,IPI可稳定操控模型输出,最有效配置在隐蔽约束下达到最高64%的攻击成功率。该研究揭示了黑盒环境下IPI的实际威胁,强调需加强针对多模态提示注入的防御。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) integrate vision and text to power applications, but this integration introduces new vulnerabilities. We study Image-based Prompt Injection (IPI), a black-box attack in which adversarial instructions are embedded into natural images to override model behavior. Our end-to-end IPI pipeline incorporates segmentation-based region selection, adaptive font scaling, and background-aware rendering to conceal prompts from human perception while preserving model interpretability. Using the COCO dataset and GPT-4-turbo, we evaluate 12 adversarial prompt strategies and multiple embedding configurations. The results show that IPI can reliably manipulate the output of the model, with the most effective configuration achieving up to 64\% attack success under stealth constraints. These findings highlight IPI as a practical threat in black-box settings and underscore the need for defenses against multimodal prompt injection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。