arXiv:2510.01494cs.LGcs.AI2025-10被引 1

攻击在数据空间可迁移,但在表示空间难以迁移,除非表示几何对齐。

Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed

  • 区分数据空间与表示空间攻击,提出迁移性取决于操作域
  • 表示空间攻击成功劫持模型但无法跨模型转移,数据空间攻击则可转移
  • 当视觉语言模型的隐空间几何对齐时,表示空间攻击也能迁移

对抗鲁棒性研究长期表明,图像分类器间的对抗样本可迁移,语言模型的越狱攻击也可迁移。然而,近期两项研究未能实现视觉语言模型间图像越狱攻击的迁移。为此,我们提出一个根本性区别:作用于输入数据空间的攻击可迁移,而作用于模型表示空间的攻击则不可迁移,至少在表示几何未对齐时如此。我们在四种不同场景中提供了理论与实证证据:首先,在两个网络计算相同输入-输出映射但使用不同表示的简化设定下,数学证明了该差异;其次,构造针对图像分类器的表示空间攻击,其成功率与经典数据空间攻击相当,但无法迁移;第三,构造针对语言模型的表示空间攻击,成功实现越狱但无法跨模型传播;第四,构造针对视觉语言模型的数据空间攻击,可成功转移到新模型,并证明当视觉语言模型的潜在几何在后投影空间中足够对齐时,表示空间攻击也可实现迁移。本工作揭示,对抗迁移并非所有攻击的固有属性,而是依赖于其操作域——共享数据空间或模型独特表示空间,这对构建更鲁棒模型具有关键意义。

原文摘要 · Abstract (English)

The field of adversarial robustness has long established that adversarial examples can successfully transfer between image classifiers and that text jailbreaks can successfully transfer between language models (LMs). However, a pair of recent studies reported being unable to successfully transfer image jailbreaks between vision-language models (VLMs). To explain this striking difference, we propose a fundamental distinction regarding the transferability of attacks against machine learning models: attacks in the input data-space can transfer, whereas attacks in model representation space do not, at least not without geometric alignment of representations. We then provide theoretical and empirical evidence of this hypothesis in four different settings. First, we mathematically prove this distinction in a simple setting where two networks compute the same input-output map but via different representations. Second, we construct representation-space attacks against image classifiers that are as successful as well-known data-space attacks, but fail to transfer. Third, we construct representation-space attacks against LMs that successfully jailbreak the attacked models but again fail to transfer. Fourth, we construct data-space attacks against VLMs that successfully transfer to new VLMs, and we show that representation space attacks can transfer when VLMs' latent geometries are sufficiently aligned in post-projector space. Our work reveals that adversarial transfer is not an inherent property of all attacks but contingent on their operational domain - the shared data-space versus models' unique representation spaces - a critical insight for building more robust models.

对抗攻击迁移性表示空间视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。