arXiv:2603.00529cs.CVcs.AI2026-03

仅修改1.2%图像块,就能让顶级图文模型生成任意目标描述。

CaptionFool: Universal Image Captioning Model Attacks

  • 通过微调7个图像块实现通用攻击,不依赖具体输入。
  • 94-96%成功率生成指定文案,包括违规内容。
  • 可生成绕过内容过滤的俚语,揭示模型安全漏洞。

图像描述模型为基于变压器的编码器-解码器结构,训练于大规模图文数据集,易受对抗攻击。本文提出CaptionFool,一种新型通用(输入无关)对抗攻击方法,针对当前最先进的基于Transformer的图像描述模型。仅需修改577个图像块中的7个(约1.2%),即可在94%-96%的条件下生成任意目标描述,包括不当内容。进一步实验表明,CaptionFool能生成专为规避现有内容审核机制设计的俚语表达。研究揭示了部署中视觉-语言模型的关键安全隐患,强调亟需构建有效防御机制。警告:本文包含具有冒犯性的模型输出。

原文摘要 · Abstract (English)

Image captioning models are encoder-decoder architectures trained on large-scale image-text datasets, making them susceptible to adversarial attacks. We present CaptionFool, a novel universal (input-agnostic) adversarial attack against state-of-the-art transformer-based captioning models. By modifying only 7 out of 577 image patches (approximately 1.2% of the image), our attack achieves 94-96% success rate in generating arbitrary target captions, including offensive content. We further demonstrate that CaptionFool can generate "slang" terms specifically designed to evade existing content moderation filters. Our findings expose critical vulnerabilities in deployed vision-language models and underscore the urgent need for robust defenses against such attacks. Warning: This paper contains model outputs which are offensive in nature.

图像生成对抗攻击安全漏洞内容过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。