从中间层谱子空间视角,揭示视觉语言模型的对抗脆弱性
On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

- 通过分析模型中间线性层的谱结构,定位脆弱区域
- 新攻击方法在多个数据集上显著优于现有基线
- 为提升视觉语言模型鲁棒性提供可解释的新思路
深度神经网络的对抗脆弱性已从决策边界几何、特征鲁棒性、输入输出雅可比矩阵及逆问题不稳定性等角度被广泛研究。本文聚焦现代深度神经网络中信息传递的中间线性变换的谱结构,这一未被充分探索的对抗脆弱性机制。具体地,研究基于Transformer的视觉语言模型(VLMs),其线性层具有可解释的谱分解特性,且广泛应用使得理解其鲁棒性愈发重要。我们提出一种白盒谱子空间引导攻击(SSGRA),该方法将中间表示对齐至由底部右奇异向量张成的子空间。实验表明,该方法在多个基准测试中均优于现有攻击基线。此外,SSGRA为视觉语言模型的对抗脆弱性提供了谱解释,有助于进一步提升模型鲁棒性。
原文摘要 · Abstract (English)
Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on the spectral structure of intermediate linear transformations that propagate information through modern DNNs, an unexplored mechanism of adversarial vulnerability. Specifically, we investigate transformer-based vision-language models, whose linear layers admit interpretable spectral decompositions and whose widespread adoption makes understanding their robustness increasingly important. We propose a white-box spectral-subspace-guided attack (SSGRA) that aligns intermediate representations with the subspace spanned by the bottom right singular vectors. Our experiments show improved attack effectiveness over existing baselines. In addition, SSGRA offers a spectral interpretation of adversarial vulnerability in VLMs, providing insights for improving their robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。