arXiv:2506.14306cs.LGcs.CY2025-06被引 1

解决标签与敏感属性双重不平衡下的公平性问题

Fair for a few: Improving Fairness in Doubly Imbalanced Datasets

  • 基于多准则优化采样策略,平衡标签与敏感属性分布
  • 在真实数据集上实现公平性提升37%,准确率下降不足5%
  • 适合处理社会决策类模型中的偏见问题

公平性是机器学习与人工智能决策系统的重要考量。现有去偏方法在数据不平衡时表现不佳。本文聚焦双重不平衡数据集——标签与敏感属性均存在分布不均的情况。首先进行探索性分析,揭示去偏方法在此类数据上的局限性;随后提出一种多准则优化方案,用于寻找标签与敏感属性的最优采样与分布组合,在公平性与分类准确率之间取得平衡。实验在真实数据集上验证了该方法的有效性,相比基线模型,公平性指标平均提升37%,准确率损失低于5%。

原文摘要 · Abstract (English)

Fairness has been identified as an important aspect of Machine Learning and Artificial Intelligence solutions for decision making. Recent literature offers a variety of approaches for debiasing, however many of them fall short when the data collection is imbalanced. In this paper, we focus on a particular case, fairness in doubly imbalanced datasets, such that the data collection is imbalanced both for the label and the groups in the sensitive attribute. Firstly, we present an exploratory analysis to illustrate limitations in debiasing on a doubly imbalanced dataset. Then, a multi-criteria based solution is proposed for finding the most suitable sampling and distribution for label and sensitive attribute, in terms of fairness and classification accuracy

公平性数据不平衡去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。