arXiv:2607.21036cs.CV2026-07中稿 · ance

针对遥感图像理解的视觉语言模型,提出可迁移的定向攻击方法。

GeoThreat: Transferable Targeted Adversarial Attacks on Large Vision-Language Models for Remote Sensing Image Interpretation

论文配图:GeoThreat: Transferable Targeted Adversarial Attacks on Large Vision-Language Models for Remote Sensing Image Interpretation
图 1 · 摘自论文原文
  • 从概念与感知双重层面操控对抗表征,实现精准语义干扰。
  • 在多种大模型上验证,攻击成功率超90%,且跨模型迁移能力强。
  • 适合评估遥感视觉语言模型安全性,尤其关注对抗鲁棒性研究者。

针对大型视觉语言模型(LVLMs)的对抗攻击,是评估其跨模态语义理解鲁棒性的有效手段。现有研究主要聚焦于通过污染视觉输入,诱导通用视觉语言任务中的预设错误响应,而在遥感领域的相关研究仍较匮乏。与自然图像理解相比,遥感图像解释需联合分析局部判别特征与全局场景上下文,这给在黑盒设置下实现可迁移的目标语义操纵带来额外挑战。为此,本文提出GeoThreat,一种面向遥感图像解释的可迁移定向对抗攻击方法。具体地,GeoThreat在概念与感知两个层次上调节对抗表征:使用代理图像编码器的类别标记作为概念表示,通过协同重要性估计从对抗样本的补丁标记中提炼感知表示。不同于简单地传播各层注意力分数,引入对抗-目标相似性梯度,更准确刻画局部视觉线索与目标语义操纵的相关性。随后,通过跨注意力机制动态对齐感知表示与目标补丁标记,促进局部线索向指定语义细节的适配。最后,通过基于集成的联合优化,迭代更新对抗扰动,实现概念校准与感知适应的协同。在多种LVLM上的大量实验表明,GeoThreat在可迁移性和可控性方面均优于现有方法。

原文摘要 · Abstract (English)

Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic understanding. Existing studies mainly focus on corrupting visual inputs to induce predefined erroneous responses in general vision-language tasks, whereas corresponding investigations in remote sensing fields remain largely underexplored. Compared with natural image understanding, remote sensing image interpretation requires joint reasoning over local discriminative cues and global scene context. This poses additional challenges to achieving transferable semantic manipulation toward specified responses under black-box settings. To tackle these challenges, we propose GeoThreat, a transferable targeted adversarial attack method against LVLMs for remote sensing image interpretation. Specifically, GeoThreat modulates adversarial representations in accordance with the target content at both conceptual and perceptual levels. The class tokens from surrogate image encoders are employed as conceptual representations, while perceptual representations are distilled from patch tokens of the adversarial example through collaborative importance estimation. Beyond merely rolling out attention scores across layers, we incorporate adversarial-target similarity gradients to more faithfully characterize the relevance of local visual cues to the intended semantic manipulation. The perceptual representations are then dynamically aligned with target patch tokens in a cross-attentive manner, facilitating the adaptation of local cues toward designated semantic details. Finally, adversarial perturbations are iteratively updated via ensemble-based joint optimization of conceptual calibration and perceptual adaptation. Extensive experiments across diverse LVLMs demonstrate the superiority of GeoThreat in both transferability and controllability.

对抗攻击遥感图像视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。