提出几何检测方法GeoDetect,识别视觉语言模型的对抗样本。
GeoDetect: Geometric Adversarial Detection for VLPs

- 基于嵌入空间几何结构,发现对抗样本偏离数据流形
- 对抗样本到随机点的平均距离显著大于正常样本
- 适用于多种攻击类型,适合提升多模态模型安全性
视觉语言预训练模型(VLPs)在实际应用中广泛使用,但易受对抗攻击。尽管单模态场景下的对抗检测方法已取得成功,但在多模态模型中的有效性仍不明确。本文研究VLP嵌入空间的几何特性,发现其具有不同于单模态视觉模型的结构性各向异性。理论分析表明,在该非均匀结构下,对抗攻击会增加干净样本与对抗样本之间的期望几何距离。具体而言,对抗样本到随机采样点的平均距离显著高于正常样本,表明对抗样本倾向于脱离数据流形。基于此,我们提出GeoDetect,通过几何得分捕捉这种离流形偏差来识别对抗样本。全面评估显示,该方法在多种VLP架构和攻击设置下均能可靠检测对抗样本,涵盖单模态、多模态及自适应攻击,为提升模型安全性和可靠性提供了稳健实用的解决方案。
原文摘要 · Abstract (English)
Vision-language pre-trained models (VLPs) are widely used in real-world applications. However, they remain vulnerable to adversarial attacks. Although adversarial detection methods have demonstrated success in single-modality settings (either vision or language), their effectiveness and reliability in multimodal models such as VLPs remain largely unexplored. In this work, we study the geometry of VLP embedding spaces and observe structured anisotropy that differs from unimodal vision models. Our theoretical analysis shows that under this anisotropic structure, adversarial attacks increase the expected geometric separation between clean and adversarial examples (AEs). Specifically, we demonstrate that AEs consistently exhibit greater expected distances to randomly sampled points than their clean counterparts, indicating that AEs tend to push representations out of manifold regions. Building on these insights, we propose GeoDetect, which leverages these off-manifold deviations via geometric scores to identify AEs. Through comprehensive evaluations, we show that our approach reliably detects AEs across diverse VLP architectures and threat settings, covering unimodal and multimodal attacks as well as adaptive attacks, thereby providing a robust and practical approach to improving the safety and reliability of these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。