arXiv:2411.02871cs.LGcs.CV2024-11被引 3

通过融合不确定性与分布建模,提升对抗训练的鲁棒性与泛化能力。

Enhancing Adversarial Robustness via Uncertainty-Aware Distributional Adversarial Training

  • 引入不确定性感知的分布对抗训练,增强对抗样本多样性。
  • 在四个数据集上实现顶尖对抗鲁棒性,同时保持自然输入性能。
  • 适合关注模型安全性的研究人员和工业部署团队。

尽管深度学习在多个领域取得显著成果,其对对抗样本的脆弱性仍是实际部署中的关键问题。对抗训练是提升模型鲁棒性的有效方法,但现有方案因依赖逐点增强策略,难以应对多样化的潜在攻击者,且对抗样本会破坏目标模型的统计信息,引入显著不确定性。为此,本文提出一种新的不确定性感知分布对抗训练方法,通过结合对抗样本的统计特性及其不确定性估计来增强对抗样本多样性。针对对抗样本与误分类干净样本对齐可能带来的负面影响,我们基于统计邻近性重构对齐参考,将对抗训练重构为干净域与对抗域之间的分布匹配框架。此外,设计了一种无需外部模型的内省梯度对齐方法,通过匹配两域输入梯度实现优化。在四个基准数据集及多种网络结构上的大量实验表明,该方法在保持自然性能的同时,实现了最先进的对抗鲁棒性。

原文摘要 · Abstract (English)

Despite remarkable achievements in deep learning across various domains, its inherent vulnerability to adversarial examples still remains a critical concern for practical deployment. Adversarial training has emerged as one of the most effective defensive techniques for improving model robustness against such malicious inputs. However, existing adversarial training schemes often lead to limited generalization ability against underlying adversaries with diversity due to their overreliance on a point-by-point augmentation strategy by mapping each clean example to its adversarial counterpart during training. In addition, adversarial examples can induce significant disruptions in the statistical information w.r.t. the target model, thereby introducing substantial uncertainty and challenges to modeling the distribution of adversarial examples. To circumvent these issues, in this paper, we propose a novel uncertainty-aware distributional adversarial training method, which enforces adversary modeling by leveraging both the statistical information of adversarial examples and its corresponding uncertainty estimation, with the goal of augmenting the diversity of adversaries. Considering the potentially negative impact induced by aligning adversaries to misclassified clean examples, we also refine the alignment reference based on the statistical proximity to clean examples during adversarial training, thereby reframing adversarial training within a distribution-to-distribution matching framework interacted between the clean and adversarial domains. Furthermore, we design an introspective gradient alignment approach via matching input gradients between these domains without introducing external models. Extensive experiments across four benchmark datasets and various network architectures demonstrate that our approach achieves state-of-the-art adversarial robustness and maintains natural performance.

对抗训练鲁棒性分布建模不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。