arXiv:2508.17680cs.LGcs.AI2025-08中稿 · presentation at EC…被引 2

提出特征空间适配器,高效提升模型抗攻击能力

Robustness Feature Adapter for Efficient Adversarial Training

  • 在特征空间设计适配器,替代传统梯度更新,降低计算开销
  • 消除鲁棒过拟合,使模型在未见攻击下仍保持高鲁棒性
  • 适用于大模型和大规模训练,适合构建可信基础模型

对抗训练(AT)结合投影梯度下降是提升模型在对抗攻击下鲁棒性的主流方法,但应用于大型主干网络时计算开销巨大,且存在鲁棒过拟合问题。本文提出一种基于适配器的新型特征空间对抗训练方法,可同时缓解上述问题。该方法通过在特征空间直接引入适配器,有效提升了内循环收敛质量,消除鲁棒过拟合现象。实验表明,该方法显著提高计算效率,并在多种主干架构及大规模对抗训练中实现更高的模型准确率,使对抗鲁棒性更好地泛化至未见过的攻击。结果验证了其在不同架构和规模下的有效性。

原文摘要 · Abstract (English)

Adversarial training (AT) with projected gradient descent is the most popular method to improve model robustness under adversarial attacks. However, computational overheads become prohibitively large when AT is applied to large backbone models. AT is also known to have the issue of robust overfitting. This paper contributes to solving both problems simultaneously towards building more trustworthy foundation models. In particular, we propose a new adapter-based approach for efficient AT directly in the feature space. We show that the proposed adapter-based approach can improve the inner-loop convergence quality by eliminating robust overfitting. As a result, it significantly increases computational efficiency and improves model accuracy by generalizing adversarial robustness to unseen attacks. We demonstrate the effectiveness of the new adapter-based approach in different backbone architectures and in AT at scale.

对抗训练特征适配大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。