arXiv:2410.17492cs.CRcs.CL2024-10EMNLP被引 5

伪装公平的后门攻击,触发后对特定群体歧视

BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers

  • 用分组条件触发器隐藏后门,正常时表现公平
  • 平均攻击成功率超85%,准确率损失极小
  • 可绕过公平性检测,适合研究安全防御者

攻击公平性至关重要,因为被破坏的模型可能引入偏差结果,在招聘、医疗和执法等敏感应用中削弱信任并加剧不平等。这凸显了理解公平机制如何被滥用的紧迫性,并推动开发兼具公平与鲁棒性的防御措施。我们提出BadFair,一种新型后门公平性攻击方法。BadFair精心构造一个在常规条件下保持准确性和公平性的模型,但一旦被特定触发器激活,就会对特定群体实施歧视并产生错误结果。此类攻击极为隐蔽危险,能规避现有公平性检测手段,维持正常使用时的公平表象。实验表明,BadFair在针对目标群体的攻击中平均成功率超过85%,仅造成微小准确率损失;且在多种数据集和模型上,均显著区分预定义的目标组与非目标组,表现出强烈歧视性。

原文摘要 · Abstract (English)

Attacking fairness is crucial because compromised models can introduce biased outcomes, undermining trust and amplifying inequalities in sensitive applications like hiring, healthcare, and law enforcement. This highlights the urgent need to understand how fairness mechanisms can be exploited and to develop defenses that ensure both fairness and robustness. We introduce BadFair, a novel backdoored fairness attack methodology. BadFair stealthily crafts a model that operates with accuracy and fairness under regular conditions but, when activated by certain triggers, discriminates and produces incorrect results for specific groups. This type of attack is particularly stealthy and dangerous, as it circumvents existing fairness detection methods, maintaining an appearance of fairness in normal use. Our findings reveal that BadFair achieves a more than 85% attack success rate in attacks aimed at target groups on average while only incurring a minimal accuracy loss. Moreover, it consistently exhibits a significant discrimination score, distinguishing between pre-defined target and non-target attacked groups across various datasets and models.

后门攻击公平性隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。