arXiv:2503.14751cs.LGcs.AI2025-03被引 3

提出可证明鲁棒的ShiftViT模型,提升视觉Transformer的抗干扰能力。

LipShiFT: A Certifiably Robust Shift-based Vision Transformer

  • 基于Lipschitz约束设计改进的ShiftViT,增强模型稳定性。
  • 在常见图像数据集上给出$ l_2 $范数下的Lipschitz常数上界。
  • 方法可扩展至大模型,推动视觉Transformer认证鲁棒性新纪录。

为基于Transformer的架构推导紧致的Lipschitz界面临重大挑战。由于输入规模大和高维注意力模块,通常成为训练过程中的关键瓶颈,导致次优结果。本研究揭示了这些方法在视觉任务中的实际限制。我们发现,基于Lipschitz的边界训练作为一种强正则化手段,同时限制了模型各层权重。聚焦于Lipschitz连续的ShiftViT模型,我们解决了在范数约束输入设置下基于Transformer架构的显著训练难题。通过在常见图像分类数据集上使用$ l_2 $范数,提供了该模型的Lipschitz常数上界估计。最终,我们证明该方法可扩展至更大模型,并在基于Transformer架构的认证鲁棒性方面达到最新水平。

原文摘要 · Abstract (English)

Deriving tight Lipschitz bounds for transformer-based architectures presents a significant challenge. The large input sizes and high-dimensional attention modules typically prove to be crucial bottlenecks during the training process and leads to sub-optimal results. Our research highlights practical constraints of these methods in vision tasks. We find that Lipschitz-based margin training acts as a strong regularizer while restricting weights in successive layers of the model. Focusing on a Lipschitz continuous variant of the ShiftViT model, we address significant training challenges for transformer-based architectures under norm-constrained input setting. We provide an upper bound estimate for the Lipschitz constants of this model using the $l_2$ norm on common image classification datasets. Ultimately, we demonstrate that our method scales to larger models and advances the state-of-the-art in certified robustness for transformer-based architectures.

视觉Transformer鲁棒性认证防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。