神经网络的特征干扰导致对抗攻击,且可预测。
Adversarial Attacks Leverage Interference Between Features in Superposition
- 利用特征超叠加导致的干扰机制解释对抗样本生成
- 合成数据中验证超叠加足以引发对抗脆弱性
- 适合关注模型内在表示与攻击可迁移性的研究者
为何存在对抗样本,且为何攻击能在不同模型间迁移?现有解释依赖高维几何、输入中的非鲁棒模式和决策边界结构,但缺乏表征层面的机制。本文表明,对抗脆弱性可能源于神经网络高效的信息编码方式。具体而言,脆弱性来自超叠加现象——网络在维度少于概念数时,被迫以非正交方式表示信息,从而产生特征干扰。这种干扰使得针对某一表征的扰动会波及其他表征,形成由干扰模式决定的脆弱性。在精确控制超叠加的合成设置中,我们证明超叠加足以引发对抗脆弱性。由此产生的攻击具有可预测性:PGD发现的扰动与基于干扰几何推导出的理论最优扰动高度对齐。在相似数据上训练的模型会产生相似的干扰模式,解释了攻击的可迁移性。进一步发现,真实图像分类器上的成功攻击也呈现出该机制预测的结构。结果表明,对抗脆弱性可能是网络表征压缩的副产品,补充了基于数据特性或架构因素的已有解释。
原文摘要 · Abstract (English)
Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decision boundary structure, but none provides a representation-level mechanism that explains why specific perturbations succeed and why attacks transfer between models. In this paper, we show that adversarial vulnerability can stem from efficient information encoding in neural networks. Specifically, vulnerability can arise from superposition - the phenomenon where networks represent more concepts than they have dimensions, forcing non-orthogonal representation and thus interference. This interference causes perturbations targeting one representation to affect others, creating vulnerabilities determined by interference patterns. In synthetic settings with precisely controlled superposition, we establish that superposition suffices to create adversarial vulnerability. The resulting attacks are predictable: PGD-discovered perturbations align with theoretically optimal perturbations derived from the interference geometry. Models trained on similar data develop similar interference patterns, explaining attack transferability. We then show that successful attacks on image classifiers exhibit the structure predicted by our proposed mechanism. These findings reveal that adversarial vulnerability can be a byproduct of networks' representational compression, complementing existing explanations based on data properties or architectural factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。