提出对称性公平鲁棒性方法,提升人脸识别等关键系统的安全与公平性。
Sy-FAR: Symmetry-based Fair Adversarial Robustness
- 以类别间攻击对称性替代绝对公平,更符合现实场景
- 在5个数据集上显著优于现有方法,提升鲁棒性与一致性
- 不仅改善攻击对称性,还缓解目标类别的隐蔽脆弱性
安全敏感的机器学习系统(如人脸识别)易受对抗样本攻击,包括现实世界中可物理实现的攻击。尽管已有多种提升对抗鲁棒性的方法,但通常导致不公平的鲁棒性:某些类别或群体更易被攻击。虽然已有技术尝试在保持鲁棒性的同时实现类别间完全公平,但多针对非关键场景。我们发现,在人脸识别等公平性敏感任务中,实现完全公平往往不可行——某些类别本就高度相似,导致相互误分类增多。因此,我们主张追求对称性:从类别i到j的攻击成功率应与从j到i相当,这更可行。直观上,类别相似性是双向关系;理论上,个体间的对称性可导出任意子组间的对称性,而其他公平性定义常难以实现组级公平。我们提出Sy-FAR,通过优化对称性、对抗鲁棒性,使用五个数据集及三种模型架构进行评估,涵盖目标型和非目标型真实攻击。结果表明,相比现有最优方法,Sy-FAR显著提升公平对抗鲁棒性;同时训练更快、结果更稳定。值得注意的是,该方法还缓解了一种新发现的不公平现象:攻击后容易被归入的目标类别,其脆弱性显著降低。
原文摘要 · Abstract (English)
Security-critical machine-learning (ML) systems, such as face-recognition systems, are susceptible to adversarial examples, including real-world physically realizable attacks. Various means to boost ML's adversarial robustness have been proposed; however, they typically induce unfair robustness: It is often easier to attack from certain classes or groups than from others. Several techniques have been developed to improve adversarial robustness while seeking perfect fairness between classes. Yet, prior work has focused on settings where security and fairness are less critical. Our insight is that achieving perfect parity in realistic fairness-critical tasks, such as face recognition, is often infeasible -- some classes may be highly similar, leading to more misclassifications between them. Instead, we suggest that seeking symmetry -- i.e., attacks from class $i$ to $j$ would be as successful as from $j$ to $i$ -- is more tractable. Intuitively, symmetry is a desirable because class resemblance is a symmetric relation in most domains. Additionally, as we prove theoretically, symmetry between individuals induces symmetry between any set of sub-groups, in contrast to other fairness notions where group-fairness is often elusive. We develop Sy-FAR, a technique to encourage symmetry while also optimizing adversarial robustness and extensively evaluate it using five datasets, with three model architectures, including against targeted and untargeted realistic attacks. The results show Sy-FAR significantly improves fair adversarial robustness compared to state-of-the-art methods. Moreover, we find that Sy-FAR is faster and more consistent across runs. Notably, Sy-FAR also ameliorates another type of unfairness we discover in this work -- target classes that adversarial examples are likely to be classified into become significantly less vulnerable after inducing symmetry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。