用自然文字骗过视觉语言模型,还能在真实世界打印生效
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
- 用大模型规划文字位置和内容,让攻击看起来像正常场景的一部分
- 在真实环境中打印后仍能骗过ChatGPT-4o等先进模型,成功率显著提升
- 无需训练,适用于安全测试与模型鲁棒性研究
大型视觉语言模型(LVLMs)在理解视觉内容方面表现卓越,但现有工作表明它们易受故意放置的对抗性文本影响。然而,这些文本通常容易被识别为异常。本文提出首个生成场景一致的字体对抗攻击的方法——SceneTAP,通过大模型代理实现视觉自然性与攻击有效性的统一。该方法采用三阶段流程:场景理解、对抗规划与无缝融合。利用链式思维推理,模型理解场景、生成有效对抗文本、规划合理位置并给出自然整合指令。随后由场景一致的TextDiffuser使用局部扩散机制执行攻击。我们还将方法拓展至真实环境,通过打印并放置生成贴纸,验证其实际效果。大量实验表明,即使对物理环境重新拍照,该对抗文本仍可成功误导ChatGPT-4o等前沿模型,攻击成功率大幅提升,同时保持视觉自然性和上下文合理性。本工作揭示了当前视觉语言模型在复杂对抗攻击下的脆弱性,并为防御机制提供新思路。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) have shown remarkable capabilities in interpreting visual content. While existing works demonstrate these models' vulnerability to deliberately placed adversarial texts, such texts are often easily identifiable as anomalous. In this paper, we present the first approach to generate scene-coherent typographic adversarial attacks that mislead advanced LVLMs while maintaining visual naturalness through the capability of the LLM-based agent. Our approach addresses three critical questions: what adversarial text to generate, where to place it within the scene, and how to integrate it seamlessly. We propose a training-free, multi-modal LLM-driven scene-coherent typographic adversarial planning (SceneTAP) that employs a three-stage process: scene understanding, adversarial planning, and seamless integration. The SceneTAP utilizes chain-of-thought reasoning to comprehend the scene, formulate effective adversarial text, strategically plan its placement, and provide detailed instructions for natural integration within the image. This is followed by a scene-coherent TextDiffuser that executes the attack using a local diffusion mechanism. We extend our method to real-world scenarios by printing and placing generated patches in physical environments, demonstrating its practical implications. Extensive experiments show that our scene-coherent adversarial text successfully misleads state-of-the-art LVLMs, including ChatGPT-4o, even after capturing new images of physical setups. Our evaluations demonstrate a significant increase in attack success rates while maintaining visual naturalness and contextual appropriateness. This work highlights vulnerabilities in current vision-language models to sophisticated, scene-coherent adversarial attacks and provides insights into potential defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。