arXiv:2510.00517cs.LGcs.CR2025-10中稿 · ICLR

差分注意力提升聚焦力却加剧对抗脆弱性,揭示选通与鲁棒性的根本权衡

Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness

  • 通过减法结构强化任务相关注意力,但引入对抗扰动下的敏感性
  • 实验显示攻击成功率更高,梯度方向频繁对立,局部敏感度显著上升
  • 适合关注模型安全与注意力机制设计的研究者,尤其在视觉-语言任务中

差分注意力(Differential Attention, DA)通过减法结构抑制冗余或噪声上下文,减少上下文幻觉,从而提升任务相关聚焦能力。然而,我们发现其结构在对抗扰动下存在固有脆弱性。理论分析表明,由减法设计引发的负梯度对齐是敏感性增强的关键原因,导致梯度范数增大和局部Lipschitz常数升高。我们在ViT/DiffViT及预训练CLIP/DiffCLIP上进行系统实验,覆盖五个数据集,结果表明:相比标准注意力,DA在对抗攻击下表现出更高的攻击成功率、频繁的梯度反向现象以及更强的局部敏感性。深度依赖实验进一步揭示:堆叠多层DA可通过深度噪声抵消缓解小扰动,但在大攻击预算下保护效果消失。总体而言,本研究揭示了核心权衡:DA虽提升干净输入下的判别聚焦能力,却牺牲了对抗鲁棒性,强调未来注意力机制需兼顾选择性与鲁棒性设计。

原文摘要 · Abstract (English)

Differential Attention (DA) has been proposed as a refinement to standard attention, suppressing redundant or noisy context through a subtractive structure and thereby reducing contextual hallucination. While this design sharpens task-relevant focus, we show that it also introduces a structural fragility under adversarial perturbations. Our theoretical analysis identifies negative gradient alignment-a configuration encouraged by DA's subtraction-as the key driver of sensitivity amplification, leading to increased gradient norms and elevated local Lipschitz constants. We empirically validate this Fragile Principle through systematic experiments on ViT/DiffViT and evaluations of pretrained CLIP/DiffCLIP, spanning five datasets in total. These results demonstrate higher attack success rates, frequent gradient opposition, and stronger local sensitivity compared to standard attention. Furthermore, depth-dependent experiments reveal a robustness crossover: stacking DA layers attenuates small perturbations via depth-dependent noise cancellation, though this protection fades under larger attack budgets. Overall, our findings uncover a fundamental trade-off: DA improves discriminative focus on clean inputs but increases adversarial vulnerability, underscoring the need to jointly design for selectivity and robustness in future attention mechanisms.

注意力机制对抗鲁棒性视觉-语言模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。