arXiv:2410.23142cs.LGcs.CV2024-10被引 6

用定向对抗训练提升模型公平性,兼顾鲁棒与公正。

FAIR-TAT: Improving Model Fairness Using Targeted Adversarial Training

  • 以定向攻击替代无目标攻击进行训练,优化类别间鲁棒性平衡。
  • 在多个数据集上实现更强的公平性,同时保持高准确率。
  • 适合关注模型公平性与鲁棒性平衡的研究者和开发者。

深度神经网络易受对抗攻击和常见干扰影响,削弱其鲁棒性。为增强模型抗干扰能力,对抗训练(AT)成为主流方法。然而,传统对抗训练常以牺牲模型公平性为代价,导致不同类别间鲁棒性差异显著:部分类别变得更强,而难检测类别则更脆弱。现有研究多聚焦于扰动图像下的公平性,却忽视了非扰动数据的准确率。此外,即使最先进的对抗训练模型在训练中表现鲁棒,面对多样化对抗威胁或常见干扰时,仍难以维持鲁棒性和公平性。本文提出一种新方法——公平定向对抗训练(FAIR-TAT),通过使用定向对抗攻击进行训练,实现了更优的对抗公平性权衡。实验证明该方法有效提升模型整体公平性与稳健性。

原文摘要 · Abstract (English)

Deep neural networks are susceptible to adversarial attacks and common corruptions, which undermine their robustness. In order to enhance model resilience against such challenges, Adversarial Training (AT) has emerged as a prominent solution. Nevertheless, adversarial robustness is often attained at the expense of model fairness during AT, i.e., disparity in class-wise robustness of the model. While distinctive classes become more robust towards such adversaries, hard to detect classes suffer. Recently, research has focused on improving model fairness specifically for perturbed images, overlooking the accuracy of the most likely non-perturbed data. Additionally, despite their robustness against the adversaries encountered during model training, state-of-the-art adversarial trained models have difficulty maintaining robustness and fairness when confronted with diverse adversarial threats or common corruptions. In this work, we address the above concerns by introducing a novel approach called Fair Targeted Adversarial Training (FAIR-TAT). We show that using targeted adversarial attacks for adversarial training (instead of untargeted attacks) can allow for more favorable trade-offs with respect to adversarial fairness. Empirical results validate the efficacy of our approach.

对抗训练模型公平性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。