用热流湍流制造对抗扰动,攻击红外遥感视觉语言模型
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

- 基于热流湍流设计通用对抗扰动,物理上更合理
- 在5个CLIP模型上平均攻击成功率48.5%,超基线10%以上
- 可使模型更自信地误判,适合安全敏感场景研究
视觉语言模型(VLMs)在安全关键的红外遥感图像中日益普及,但其对抗鲁棒性尚未被检验。我们提出AirflowAttack,据知是首个针对红外遥感VLM的对抗攻击,也是首个将热流湍流作为扰动先验的攻击方法。轻量级生成器合成单一输入无关扰动,正则化为物理上合理的气流模式。在单个替代CLIP模型上优化后,在五个不同的CLIP骨干网络上实现平均零样本场景分类攻击成功率(ASR)48.5%,显著高于四种红外特定物理基线(27.7%–37.0%)。应用于六个先进VLM时,最高使场景分类准确率下降38.2%相对值,却意外让部分模型对其红外分析更加自信,将扰动误认为真实的热证据(如温度梯度、对流)。消融实验表明,气流先验提升物理合理性,且未明显影响攻击成功率。结合涵盖十一个模型和四个任务的基准测试,这些发现揭示了快速扩展的红外VLM生态系统的严重漏洞。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustness remains unexamined. We present AirflowAttack, to our knowledge the first adversarial attack for IR remote-sensing VLMs and the first to weaponize thermal-airflow turbulence as the perturbation prior. A lightweight generator synthesizes a single input-agnostic perturbation regularized toward physically plausible airflow patterns. Optimized on one surrogate CLIP model, it attains a mean zero-shot scene-classification attack success rate (ASR, the fraction of samples whose top-1 class changes) of 48.5% across five diverse CLIP backbones, far exceeding four IR-specific physical baselines (27.7--37.0%). Applied to six state-of-the-art VLMs, it cuts scene-classification accuracy by up to 38.2% relative, yet paradoxically makes some models more confident in their IR analysis, confabulating the perturbation as genuine thermal evidence such as temperature gradients and convection. Ablations show the airflow prior raises physical plausibility at no measurable cost to attack success. Together with a benchmark spanning eleven models and four tasks, these findings expose critical vulnerabilities in the rapidly expanding IR VLM ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。