提出可物理部署的通用干扰贴片,破坏红外视觉语言模型多任务性能
Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models

- 设计低频曲面网格贴片,通过表征驱动优化实现跨任务干扰
- 单个贴片在多种红外模型上同时降低分类、描述与问答准确率
- 具备真实场景物理有效性,适用于安全评估与系统加固研究
红外视觉语言模型(IR-VLMs)在低可见环境下日益重要,但其对物理语义攻击的鲁棒性仍不明确。现有对抗贴片方法主要针对RGB图像或封闭集红外检测器,无法直接应对开放式的IR-VLM任务——即一个可部署实体能同时影响分类、生成描述和视觉问答。本文提出通用曲面网格贴片(UCGP),将贴片表示为可部署的低频曲面网格,通过子空间偏离、拓扑破坏和局部外观正则化联合优化。采用元差分进化(MetaDE)搜索贴片参数,并引入期望变换(EOT)采样以应对成像变化,以及薄板样条(TPS)建模非刚性形变。不同于操纵标签或提示,UCGP通过破坏视觉表征空间中的干净类别流形,导致跨模态输出质量下降。实验表明,单一共享贴片在多种IR-VLM架构上同时削弱分类、生成与问答性能,且具有显著跨模型迁移性、跨数据集泛化性、跨类别扩展性及真实场景物理有效性。该结果揭示当前红外多模态系统存在鲁棒性盲区。代码与复现资源见https://github.com/dyx6663/UCGP。
原文摘要 · Abstract (English)
Infrared vision-language models (IR-VLMs) are becoming important for semantic perception in low-visibility environments, yet their robustness to physical semantic attacks remains underexplored. Existing adversarial patch methods are mainly designed for red-green-blue (RGB) images or closed-set infrared detectors and do not directly address open-ended IR-VLM tasks, where one deployable artifact can affect classification, captioning, and visual question answering (VQA) simultaneously. We propose Universal Curved-Grid Patch, abbreviated UCGP, a universal physical adversarial patch framework tailored to IR-VLMs. UCGP represents the patch as a deployable low-frequency curved grid and optimizes it with a representation-driven objective over subspace departure, topology disruption, and local appearance regularization. Meta Differential Evolution (MetaDE) searches for patch parameters with two physical robustness augmentations: Expectation over Transformation (EOT) capture sampling for imaging variations and thin-plate spline (TPS) patch-deformation modeling for non-rigid local shape changes. Rather than manipulating labels or prompts, UCGP disrupts the clean-category manifold in visual representation space, which later appears as degraded cross-modal outputs. Experiments show that a single shared patch degrades classification, captioning, and VQA across diverse IR-VLM architectures while retaining measurable cross-model transfer, cross-dataset generalization, cross-category extensibility, and real-scene physical effectiveness under the evaluated conditions. These results reveal a robustness blind spot in current infrared multimodal systems. Code and reproducibility assets are available at https://github.com/dyx6663/UCGP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。