研究文字图像攻击与模型对齐关系,揭示字体大小和视觉干扰如何影响攻击成功率。
Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

- 通过分析文本渲染为图像的攻击方式,发现字体大小显著影响攻击效果。
- 嵌入距离越小,攻击成功率越高,相关性达-0.93(p<0.01)。
- 不同模型对视觉干扰响应差异大,无通用防御方案。
我们研究了针对视觉语言模型(VLM)的字形提示注入攻击,其中恶意文本被渲染为图像以绕过安全机制。此类攻击威胁日益增长,因VLM正作为自主代理的感知核心,应用于浏览器自动化、计算机使用系统及带摄像头的具身智能体。实际攻击面异质性强:恶意文本在不同字体大小和多样视觉条件下呈现,且各类VLM的脆弱性差异显著,导致防御复杂化。我们在四个VLM(GPT-4o、Claude Sonnet 4.5、Mistral-Large-3、Qwen3-VL-4B-Instruct)上评估了来自SALAD-Bench的1000个提示,在6–28像素字体大小及旋转、模糊、噪声、对比度变化等视觉变换下发现:(1) 字体大小显著影响攻击成功率(ASR),极小字体(6px)几乎无效,中等字体最有效;(2) 对GPT-4o(36% vs 8%)和Claude(47% vs 22%),文本攻击优于图像攻击,而Qwen3-VL与Mistral在模态间表现相近;(3) 两种多模态嵌入模型(JinaCLIP与Qwen3-VL-Embedding)的文本-图像嵌入距离与所有四模型的ASR呈强负相关(r = -0.71至-0.93,p < 0.01);(4) 严重退化使嵌入距离增加10–12%,并导致ASR下降34–96%;旋转对模型影响不对称(Mistral下降50%,GPT-4o不变)。这些结果表明,模型特异性鲁棒性模式排除了通用防御策略,并为在对抗环境中部署智能体的实践者选择合适的VLM后端提供实证指导。
原文摘要 · Abstract (English)
We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms, posing a growing threat as VLMs serve as the perceptual backbone of autonomous agents, from browser automation and computer-use systems to camera-equipped embodied agents. In practice, the attack surface is heterogeneous: adversarial text appears at varying font sizes and under diverse visual conditions, while the growing ecosystem of VLMs exhibits substantial variation in vulnerability, complicating defensive approaches. Evaluating 1,000 prompts from SALAD-Bench across four VLMs, namely, GPT-4o, Claude Sonnet 4.5, Mistral-Large-3, and Qwen3-VL-4B-Instruct under varying font sizes (6--28px) and visual transformations (rotation, blur, noise, contrast changes), we find: (1) font size significantly affects attack success rate (ASR), with very small fonts (6px) yielding near-zero ASR while mid-range fonts achieve peak effectiveness; (2) text attacks are more effective than image attacks for GPT-4o (36% vs 8%) and Claude (47% vs 22%), while Qwen3-VL and Mistral show comparable ASR across modalities; (3) text-image embedding distance from two multimodal embedding models (JinaCLIP and Qwen3-VL-Embedding) shows strong negative correlation with ASR across all four models (r = -0.71 to -0.93, p < 0.01); (4) heavy degradations increase embedding distance by 10--12% and reduce ASR by 34--96%, while rotation asymmetrically affects models (Mistral drops 50%, GPT-4o unchanged). These findings highlight that model-specific robustness patterns preclude one-size-fits-all defenses and offer empirical guidance for practitioners selecting VLM backbones for agentic systems operating in adversarial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。