小灰块可骗过红外视觉模型,让其输出指定错误结果。
InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models

- 用小块灰图精准攻击红外视觉语言模型,控制扰动范围在5%内。
- 在分类、描述、问答任务中成功率达86%至100%。
- 揭示不同模型对攻击的敏感差异,适合安全评估与防御研究者参考。
红外视觉语言模型(IR-VLMs)在低可见环境下表现出色,但其对定向对抗攻击的鲁棒性仍不清晰。现有方法多针对彩色图像模型或单一任务,未验证局部扰动能否在IR-VLM中引发特定语义目标。本文提出InfraPatch,一种白盒、实例级的数字灰度补丁攻击框架,可在约5%局部区域预算内优化单通道补丁,结合代理引导定位与任务自适应语义目标,实现图像分类、图像描述和二元视觉问答中的目标行为诱导。在300张通过DiffV2IR生成的合成红外风格图像上评估10种红外适配模型变体,采用清洁条件下的定向成功标准。InfraPatch在各变体中实现86.00%至100%的攻击成功率。相较优化随机放置,代理位置搜索分别提升CLIP与BLIP-2成功率6.67和10.33个百分点;LLaVA-1.5在两种设置下均接近100%饱和。补丁面积与目标函数消融实验进一步暴露不同架构与任务格式间的显著脆弱性差异。结果表明,在可控数字威胁模型下,小灰块可跨模型家族注入指定语义,推动红外多模态系统更强鲁棒性评估。
原文摘要 · Abstract (English)
Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception under low-visibility conditions, yet their robustness to targeted adversarial attacks remains poorly understood. Existing adversarial patch methods mainly study RGB-based models or a single downstream task and do not characterize whether localized perturbations can induce an intended semantic target in IR-VLMs. We propose InfraPatch, a white-box, per-instance framework for targeted digital grayscale patch attacks against IR-VLMs. InfraPatch optimizes a compact single-channel patch within an approximately 5% local-area budget, combines proxy-guided placement with task-adaptive semantic objectives, and induces target behaviors in image classification, image captioning, and binary visual question answering. We evaluate ten infrared-adapted model variants on 300 synthetic infrared-style images generated by applying DiffV2IR to a fixed 30-category COCO subset, using clean-conditioned targeted success criteria. InfraPatch achieves targeted attack success rates from 86.00% to 100% across the ten variants. On CLIP and BLIP-2, proxy location search improves success by 6.67 and 10.33 percentage points over optimized random placement, respectively; LLaVA-1.5 remains saturated near 100% under both settings. Patch-area and objective ablations further expose substantial differences in vulnerability across architectures and task formats. These results show that small grayscale patches can inject chosen target semantics across IR-VLM families under a controlled digital threat model, motivating stronger robustness evaluation for infrared multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。