arXiv:2501.04527cs.LGcs.CV2025-01被引 3

提出新方法提升模型对不同类别攻击的公平鲁棒性。

Towards Fair Class-wise Robustness: Class Optimal Distribution Adversarial Training

  • 构建分布鲁棒优化框架,联合优化类别权重与模型参数。
  • 理论推导出最优权重闭式解,实现全局最优。
  • 引入公平弹性系数评估,适合关注鲁棒公平性的研究者。

对抗训练虽能有效提升深度神经网络对对抗攻击的鲁棒性,但存在类别间鲁棒性差异大的公平性问题。现有基于类别重加权的方法缺乏理论支撑,且权重优化与模型参数解耦,导致方向不一致,可能产生次优结果。本文提出新型极小极大训练框架——类最优分布对抗训练(CODAT),利用分布鲁棒优化充分探索类别权重空间,理论上保证最优权重。通过推导内层最大化问题的闭式解,获得确定性等价目标函数,支持权重与模型参数的联合优化。同时提出公平弹性系数用于评估算法在鲁棒性与鲁棒公平性上的表现。在多个数据集上的实验表明,该方法显著提升模型鲁棒公平性,优于当前最先进方法。

原文摘要 · Abstract (English)

Adversarial training has proven to be a highly effective method for improving the robustness of deep neural networks against adversarial attacks. Nonetheless, it has been observed to exhibit a limitation in terms of robust fairness, characterized by a significant disparity in robustness across different classes. Recent efforts to mitigate this problem have turned to class-wise reweighted methods. However, these methods suffer from a lack of rigorous theoretical analysis and are limited in their exploration of the weight space, as they mainly rely on existing heuristic algorithms or intuition to compute weights. In addition, these methods fail to guarantee the consistency of the optimization direction due to the decoupled optimization of weights and the model parameters. They potentially lead to suboptimal weight assignments and consequently, a suboptimal model. To address these problems, this paper proposes a novel min-max training framework, Class Optimal Distribution Adversarial Training (CODAT), which employs distributionally robust optimization to fully explore the class-wise weight space, thus enabling the identification of the optimal weight with theoretical guarantees. Furthermore, we derive a closed-form optimal solution to the internal maximization and then get a deterministic equivalent objective function, which provides a theoretical basis for the joint optimization of weights and model parameters. Meanwhile, we propose a fairness elasticity coefficient for the evaluation of the algorithm with regard to both robustness and robust fairness. Experimental results on various datasets show that the proposed method can effectively improve the robust fairness of the model and outperform the state-of-the-art approaches.

对抗训练鲁棒性公平性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。