arXiv:2607.07288cs.CV2026-07

提出一种贴边放置的二维码结构化攻击,让红外视觉语言模型失效

InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models

论文配图:InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models
图 1 · 摘自论文原文
  • 在图像边缘放置可学习的结构化补丁,模仿二维码形态
  • 使CLIP类模型准确率从98.67%降至0.70%,跨任务攻击有效
  • 适用于研究红外视觉模型安全性的研究人员

红外视觉语言模型在低光和恶劣视觉条件下日益广泛应用,但其对局部结构化扰动的鲁棒性仍缺乏系统研究。现有红外对抗研究主要聚焦目标检测器,对红外视觉语言模型的安全性关注不足。本文提出InfraQR,一种受二维码启发的结构化补丁攻击方法,针对红外视觉语言模型。与将扰动附加到目标物体的局部攻击不同,InfraQR将紧凑的结构化补丁置于图像边界,并通过替代的CLIP风格编码器优化可学习网格单元。生成的补丁具有近似二值的结构化外观,但无需是合法或机器可读的二维码。我们在300张红外图像基准上评估了InfraQR在红外分类、图文生成迁移和问答感知视觉问答(VQA)任务上的表现。结果表明,InfraQR显著降低多个CLIP风格分类器的准确率,包括将OpenAI CLIP准确率从98.67%降至0.70%。生成的对抗图像还能迁移至黑盒图文生成和VQA模型,导致描述语义退化,并在基于GPT-5.4的评估下产生更多错误回答。这些结果表明,红外视觉语言模型仍易受结构化边缘扰动影响,提示需进一步研究超越直接目标遮挡的跨任务鲁棒性。

原文摘要 · Abstract (English)

Infrared vision-language models are increasingly used for perception under low-light and adverse visual conditions, yet their robustness to localized structured perturbations remains underexplored. Existing infrared adversarial studies mainly focus on object detectors, leaving the security of infrared vision-language models less systematically examined. We present InfraQR, a QR-inspired structured patch attack for infrared vision-language models. Unlike localized attacks that attach perturbations to the target object, InfraQR places a compact structured patch along image boundaries and optimizes learnable grid cells through surrogate CLIP-style encoders. The resulting patch has a near-binary structured appearance, but is not required to be a valid or machine-readable QR code. We evaluate InfraQR on infrared classification, caption transfer, and question-answer-aware visual question answering (VQA) tasks. On a 300-image infrared benchmark, InfraQR sharply reduces the accuracy of multiple CLIP-style classifiers, including reducing OpenAI CLIP accuracy from 98.67% to 0.70%. The generated adversarial images also transfer to black-box captioning and VQA models, causing semantic degradation in captions and more error-prone answers under GPT-5.4-based evaluation. These results show that infrared vision-language models remain vulnerable to structured edge-placed perturbations, motivating further study of cross-task robustness beyond direct object occlusion.

红外视觉对抗攻击结构化扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。