arXiv:2411.14424cs.LGcs.CR2024-11被引 3

用混合数据提升分类器对各类别的公平鲁棒性

Learning Fair Robustness via Domain Mixup

  • 同类别样本混合后做对抗训练,均衡各分类鲁棒性
  • 理论证明可降低类别间对抗风险差异,自然风险也更均等
  • 适合关注模型公平性与鲁棒性的研究者

对抗训练是提升分类器对抗攻击鲁棒性的主流方法,但其往往无法保证各类别获得同等的鲁棒性。本文提出利用混合(mixup)策略来学习公平的鲁棒分类器,通过混合同类别样本并在混合样本上进行对抗训练,实现跨类别的鲁棒性均衡。针对线性分类器,我们给出了理论分析,证明混合与对抗训练结合可严格减少类别间的鲁棒性差异。该方法不仅降低类别间对抗风险差异,还改善了自然风险的分布。实验在合成数据和真实数据集CIFAR-10上验证了其有效性,显著减小了自然风险与对抗风险在各类别间的差距。

原文摘要 · Abstract (English)

Adversarial training is one of the predominant techniques for training classifiers that are robust to adversarial attacks. Recent work, however has found that adversarial training, which makes the overall classifier robust, it does not necessarily provide equal amount of robustness for all classes. In this paper, we propose the use of mixup for the problem of learning fair robust classifiers, which can provide similar robustness across all classes. Specifically, the idea is to mix inputs from the same classes and perform adversarial training on mixed up inputs. We present a theoretical analysis of this idea for the case of linear classifiers and show that mixup combined with adversarial training can provably reduce the class-wise robustness disparity. This method not only contributes to reducing the disparity in class-wise adversarial risk, but also the class-wise natural risk. Complementing our theoretical analysis, we also provide experimental results on both synthetic data and the real world dataset (CIFAR-10), which shows improvement in class wise disparities for both natural and adversarial risks.

对抗训练公平性混合数据鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。