arXiv:2503.06150cs.LG2025-03中稿 · IEEE Transactions …被引 3

公平性干预未必损害隐私,反而增强抗攻击能力。

Do Fairness Interventions Come at the Cost of Privacy: Evaluations for Binary Classifiers

  • 通过去除特征中的敏感信息降低模型置信度,提升隐私防护。
  • 公平模型对成员推断和属性推断攻击的抵抗力更强。
  • 发现公平与非公平模型预测差异可被用于高级攻击,暴露新风险。

尽管事前公平性方法在缓解偏差预测方面展现出潜力,但其对隐私泄露的影响仍缺乏深入研究。本文通过成员推断攻击(MIAs)和属性推断攻击(AIAs)评估公平增强型二分类器的隐私风险。令人意外的是,公平性干预并未导致隐私妥协,反而提升了对两类攻击的防御能力。这是因为公平干预通常会移除提取特征中的敏感信息,并降低多数训练数据的置信度以实现更公平的预测。然而,在评估中我们发现一种潜在威胁机制:利用公平模型与非公平模型之间的预测差异,可显著提升MIAs和AIAs的攻击效果。该机制揭示了公平模型的强脆弱性,对当前公平方法构成重大隐私风险。在多个数据集、攻击方法和代表性公平策略上的广泛实验验证了上述发现,并证明了该机制的有效性。本研究揭示了公平性研究中被忽视的隐私威胁,呼吁在模型部署前进行全面的安全漏洞评估。

原文摘要 · Abstract (English)

While in-processing fairness approaches show promise in mitigating biased predictions, their potential impact on privacy leakage remains under-explored. We aim to address this gap by assessing the privacy risks of fairness-enhanced binary classifiers via membership inference attacks (MIAs) and attribute inference attacks (AIAs). Surprisingly, our results reveal that enhancing fairness does not necessarily lead to privacy compromises. For example, these fairness interventions exhibit increased resilience against MIAs and AIAs. This is because fairness interventions tend to remove sensitive information among extracted features and reduce confidence scores for the majority of training data for fairer predictions. However, during the evaluations, we uncover a potential threat mechanism that exploits prediction discrepancies between fair and biased models, leading to advanced attack results for both MIAs and AIAs. This mechanism reveals potent vulnerabilities of fair models and poses significant privacy risks of current fairness methods. Extensive experiments across multiple datasets, attack methods, and representative fairness approaches confirm our findings and demonstrate the efficacy of the uncovered mechanism. Our study exposes the under-explored privacy threats in fairness studies, advocating for thorough evaluations of potential security vulnerabilities before model deployments.

公平性隐私保护攻击检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。