arXiv:2502.21041cs.LGcs.AI2025-02被引 2

解决稀疏攻击下快速对抗训练的过拟合问题

Fast Adversarial Training against Sparse Attacks Requires Loss Smoothing

  • 引入软标签与权衡损失函数平滑损失曲面
  • 1步攻击下实现媲美多步攻击的性能
  • 适合防御稀疏扰动的高效模型训练

本文研究针对 $l_0$ 范数约束的稀疏对抗扰动的快速对抗训练。实验表明,采用单步攻击在 $l_0$ 限制扰动上会导致性能下降和灾难性过拟合(CO)。分析揭示,$l_0$ 对抗训练中的 CO 源于单步攻击产生的次优扰动位置。理论与实证分析显示,$l_0$ 对抗训练的损失曲面比 $l_ ext{∞}$、$l_2$ 和 $l_1$ 更加崎岖,且崎岖曲面会加剧 CO。为此,提出 Fast-LS-$l_0$,通过引入软标签与权衡损失函数来平滑对抗损失曲面。大量实验表明,该方法可有效克服灾难性过拟合,达到当前最优性能,并缩小单步与多步对抗训练在稀疏攻击下的性能差距。

原文摘要 · Abstract (English)

This paper studies fast adversarial training against sparse adversarial perturbations bounded by $l_0$ norm. We demonstrate the challenges of employing $1$-step attacks on $l_0$ bounded perturbations for fast adversarial training, including degraded performance and the occurrence of catastrophic overfitting (CO). We highlight that CO in $l_0$ adversarial training is caused by sub-optimal perturbation locations of $1$-step attack. Theoretical and empirical analyses reveal that the loss landscape of $l_0$ adversarial training is more craggy compared to its $l_\infty$, $l_2$ and $l_1$ counterparts. Moreover, we corroborate that the craggy loss landscape can aggravate CO. To address these issues, we propose Fast-LS-$l_0$ that incorporates soft labels and the trade-off loss function to smooth the adversarial loss landscape. Extensive experiments demonstrate our method can overcome the challenge of catastrophic overfitting, achieve state-of-the-art performance, and narrow down the performance gap between $1$-step and multi-step adversarial training against sparse attacks.

对抗训练稀疏攻击损失平滑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。