通过嵌入文字图像实现对多模态智能体的自适应攻击
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- 用优化文字嵌入网页图,隐蔽诱导多模态模型执行恶意指令
- 在仅用图像的攻击下成功率从23%提升至45%,跨多个模型有效
- 支持持续学习,可复用历史成功策略增强后续攻击
基于大视觉语言模型的多模态智能体在开放环境中日益普及,但极易受到提示注入攻击,尤其是通过视觉输入。我们提出AgentTypo,一个黑盒红队框架,通过将优化文本嵌入网页图像实现自适应排版提示注入。其自动排版提示注入(ATPI)算法通过替换字幕器最大化提示重建,同时通过隐蔽损失最小化人类察觉,利用树状帕尔策估计器在文本位置、大小和颜色上进行黑盒优化。为增强攻击能力,我们开发AgentTypo-pro,一个多大语言模型系统,通过评估反馈迭代优化注入提示,并检索过往成功案例实现持续学习。有效提示被抽象为通用策略并存入策略库,实现知识积累与复用。在VWA-Adv基准测试中,涵盖分类广告、购物和Reddit场景,AgentTypo显著优于最新图像攻击方法如AgentAttack。在GPT-4o智能体上,仅图像攻击使成功率从0.23升至0.45,且在GPT-4V、GPT-4o-mini、Gemini 1.5 Pro和Claude 3 Opus上均表现一致。在图像+文本场景下,实现0.68的攻击成功率,同样超越现有基线。结果表明,AgentTypo对多模态智能体构成现实而强大的威胁,凸显防御机制的紧迫性。
原文摘要 · Abstract (English)
Multimodal agents built on large vision-language models (LVLMs) are increasingly deployed in open-world settings but remain highly vulnerable to prompt injection, especially through visual inputs. We introduce AgentTypo, a black-box red-teaming framework that mounts adaptive typographic prompt injection by embedding optimized text into webpage images. Our automatic typographic prompt injection (ATPI) algorithm maximizes prompt reconstruction by substituting captioners while minimizing human detectability via a stealth loss, with a Tree-structured Parzen Estimator guiding black-box optimization over text placement, size, and color. To further enhance attack strength, we develop AgentTypo-pro, a multi-LLM system that iteratively refines injection prompts using evaluation feedback and retrieves successful past examples for continual learning. Effective prompts are abstracted into generalizable strategies and stored in a strategy repository, enabling progressive knowledge accumulation and reuse in future attacks. Experiments on the VWA-Adv benchmark across Classifieds, Shopping, and Reddit scenarios show that AgentTypo significantly outperforms the latest image-based attacks such as AgentAttack. On GPT-4o agents, our image-only attack raises the success rate from 0.23 to 0.45, with consistent results across GPT-4V, GPT-4o-mini, Gemini 1.5 Pro, and Claude 3 Opus. In image+text settings, AgentTypo achieves 0.68 ASR, also outperforming the latest baselines. Our findings reveal that AgentTypo poses a practical and potent threat to multimodal agents and highlight the urgent need for effective defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。