arXiv:2511.08331cs.LG2025-11

通过少量恶意数据注入,可严重破坏分类模型的公平性。

Adversarial Bias: Data Poisoning Attacks on Fairness

  • 向训练集注入精心设计的对抗数据点,改变模型决策边界。
  • 在多个数据集上显著降低公平性指标,准确率影响微小。
  • 方法通用性强,对多种模型均有效,适用于研究公平性漏洞者。

随着人工智能和机器学习系统在现实应用中的普及,确保其公平性变得日益关键。现有研究多聚焦于评估与提升模型公平性,但对公平性脆弱性的研究较少,即如何被故意破坏。本文首次从理论上证明,一种简单的对抗性数据投毒策略足以使朴素贝叶斯分类器产生最大不公平行为。核心思路是:在训练集中巧妙注入少量精心构造的对抗样本,使模型决策边界偏向特定保护群体,同时保持整体性能。实验表明,该攻击在多个基准数据集和模型上均显著优于现有方法,大幅降低公平性指标,而准确率损失极小。尤其值得注意的是,该方法对多种模型具有普适性,展现了对机器学习系统公平性的强大且稳健的威胁能力。

原文摘要 · Abstract (English)

With the growing adoption of AI and machine learning systems in real-world applications, ensuring their fairness has become increasingly critical. The majority of the work in algorithmic fairness focus on assessing and improving the fairness of machine learning systems. There is relatively little research on fairness vulnerability, i.e., how an AI system's fairness can be intentionally compromised. In this work, we first provide a theoretical analysis demonstrating that a simple adversarial poisoning strategy is sufficient to induce maximally unfair behavior in naive Bayes classifiers. Our key idea is to strategically inject a small fraction of carefully crafted adversarial data points into the training set, biasing the model's decision boundary to disproportionately affect a protected group while preserving generalizable performance. To illustrate the practical effectiveness of our method, we conduct experiments across several benchmark datasets and models. We find that our attack significantly outperforms existing methods in degrading fairness metrics across multiple models and datasets, often achieving substantially higher levels of unfairness with a comparable or only slightly worse impact on accuracy. Notably, our method proves effective on a wide range of models, in contrast to prior work, demonstrating a robust and potent approach to compromising the fairness of machine learning systems.

公平性数据投毒对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。