arXiv:2607.29445cs.CVcs.AI2026-07

用二维码热图诱骗红外视觉模型误判语义,隐蔽且无需训练。

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

论文配图:QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
图 1 · 摘自论文原文
  • 基于二维码结构设计可调温区的热扰动模式。
  • 在多个任务上成功引导模型输出目标语义,保持视觉不可见。
  • 适合研究红外模型安全或对抗攻击的学者使用。

红外视觉语言模型(IR-VLMs)将热成像感知拓展至开放词汇分类、图像描述和视觉问答等任务,但其对结构化热扰动的鲁棒性及跨模态语义对齐稳定性仍不足。本文提出无训练、黑盒的QR-Structured Thermal Triggers(QR-STT),通过优化二维码内部模块的冷、中、热状态分布,联合搜索位置、尺度、旋转、强度、模糊度和圆润度等参数。采用三阶段无梯度搜索与贪婪模块翻转优化,实现离散与连续空间的高效协同。目标函数促进与攻击者选定目标的对齐,抑制源类别证据,并正则化二维码结构与视觉相似性。在多个CLIP-style编码器上实验表明,QR-STT能持续引导图像-文本对齐至指定概念,同时保持视觉隐蔽性。针对分类优化的扰动可迁移至图像描述与视觉问答,引发生成结果的一致语义漂移。结果揭示二维码结构热模式是语言驱动红外感知的可解释攻击面,凸显了对跨任务结构化语义攻击进行鲁棒性评估的必要性。

原文摘要 · Abstract (English)

Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while optimizing its internal modules, each of which is assigned a cold, neutral, or hot thermal state. The framework jointly searches module topology and rendering parameters, including position, scale, rotation, intensity, blur, and roundness. A three-stage gradient-free procedure with greedy module-flip refinement efficiently handles the mixed discrete and continuous search space. The objective promotes alignment with an attacker-selected target, suppresses source-class evidence, and regularizes QR structure and visual similarity. Experiments on multiple CLIP-style encoders show that QR-STT consistently redirects image-text alignment toward chosen concepts while maintaining visual stealth. Perturbations optimized for classification also transfer to image captioning and VQA, causing target-consistent semantic drift in generated outputs. These results identify QR-structured thermal patterns as an interpretable attack surface for language-driven infrared perception and highlight the need for robustness evaluation against structured cross-task semantic attacks.

对抗攻击红外感知视觉语言模型热扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。