arXiv:2412.19947cs.LGcs.AI2024-12被引 2

通过改进标准差机制提升模型对抗攻击的鲁棒性与泛化能力

Standard-Deviation-Inspired Regularization for Improving Adversarial Robustness

  • 基于输出概率的标准差设计正则化项,增强模型对对抗样本的敏感性
  • 在CW和Auto-attack下显著提升鲁棒性,准确率提升超过5%以上
  • 适用于需要高鲁棒性的实际场景,如安全关键系统

对抗训练(AT)已被证明能有效提升深度神经网络(DNN)对对抗攻击的鲁棒性。AT是一种极小极大优化过程,其中内层最大化步骤生成对抗样本以增强模型的鲁棒性,外层最小化则降低这些对抗样本上的损失。本文提出一种受标准差启发的(SDI)正则化项,用于进一步提升模型鲁棒性和泛化性能。我们指出,AT中的内层最大化类似于最小化模型输出概率的修正标准差;而最大化的这一修正标准差可与外层最小化互补。实验表明,该SDI度量可用于构造对抗样本,并且将SDI正则化与现有AT变体结合后,能显著提升DNN在更强攻击(如CW和Auto-attack)下的鲁棒性,同时改善泛化能力。

原文摘要 · Abstract (English)

Adversarial Training (AT) has been demonstrated to improve the robustness of deep neural networks (DNNs) against adversarial attacks. AT is a min-max optimization procedure where in adversarial examples are generated to train a more robust DNN. The inner maximization step of AT increases the losses of inputs with respect to their actual classes. The outer minimization involves minimizing the losses on the adversarial examples obtained from the inner maximization. This work proposes a standard-deviation-inspired (SDI) regularization term to improve adversarial robustness and generalization. We argue that the inner maximization in AT is similar to minimizing a modified standard deviation of the model's output probabilities. Moreover, we suggest that maximizing this modified standard deviation can complement the outer minimization of the AT framework. To support our argument, we experimentally show that the SDI measure can be used to craft adversarial examples. Additionally, we demonstrate that combining the SDI regularization term with existing AT variants enhances the robustness of DNNs against stronger attacks, such as CW and Auto-attack, and improves generalization.

对抗训练正则化鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。