arXiv:2604.19724cs.LGcs.AI2026-04被引 1

首次证明ViT在对抗训练下可实现无害过拟合,提升鲁棒泛化能力。

Benign Overfitting in Adversarial Training for Vision Transformers

论文配图:Benign Overfitting in Adversarial Training for Vision Transformers
图 1 · 摘自论文原文
  • 在特定信噪比与扰动范围内,对抗训练使ViT接近零鲁棒训练误差。
  • 即使存在过拟合现象,模型仍能保持强鲁棒泛化性能。
  • 适用于关注模型鲁棒性与理论分析的研究者,尤其针对ViT架构。

尽管视觉变换器(ViTs)在众多视觉任务中表现卓越,但近期研究发现其仍易受对抗样本影响,与卷积神经网络(CNNs)类似。对抗训练是常见的经验防御策略,但其在ViTs中的理论基础尚不明确。本文首次对简化版ViT架构下的对抗训练进行理论分析。结果表明,在满足特定信噪比条件且扰动预算适中时,对抗训练可使ViTs在某些条件下实现近乎零的鲁棒训练损失与鲁棒泛化误差。令人惊讶的是,这种情况下仍能实现强泛化能力,即所谓的‘无害过拟合’——此前仅在经对抗训练的CNN中观察到。在合成数据与真实数据集上的实验进一步验证了上述理论结论。

原文摘要 · Abstract (English)

Despite the remarkable success of Vision Transformers (ViTs) across a wide range of vision tasks, recent studies have revealed that they remain vulnerable to adversarial examples, much like Convolutional Neural Networks (CNNs). A common empirical defense strategy is adversarial training, yet the theoretical underpinnings of its robustness in ViTs remain largely unexplored. In this work, we present the first theoretical analysis of adversarial training under simplified ViT architectures. We show that, when trained under a signal-to-noise ratio that satisfies a certain condition and within a moderate perturbation budget, adversarial training enables ViTs to achieve nearly zero robust training loss and robust generalization error under certain regimes. Remarkably, this leads to strong generalization even in the presence of overfitting, a phenomenon known as \emph{benign overfitting}, previously only observed in CNNs (with adversarial training). Experiments on both synthetic and real-world datasets further validate our theoretical findings.

视觉变换器对抗训练过拟合理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。