通过近邻关系增强数据混合,提升模型公平性。
ProxiMix: Enhancing Fairness with Proximity Samples in Subgroups
- 利用近邻信息改进数据混合,生成更公平的标签。
- 在三个数据集上验证,显著改善预测与救济公平性。
- 适合关注模型公平性的研究者和实践者。
许多偏差缓解方法被提出以解决机器学习中的公平性问题。我们发现,仅使用线性混合(mixup)这一数据增强技术进行偏差缓解时,仍可能保留数据标签中的偏见。本文旨在解决该问题,提出一种新的预处理策略,结合现有mixup方法与新提出的偏差缓解算法,生成更具近邻感知能力的增强样本标签。具体地,我们提出了ProxiMix,其同时保持成对关系与近邻关系,实现更公平的数据增强。我们在三个数据集、三种机器学习模型及不同超参数设置下进行了全面实验。结果表明,从预测公平性和救济公平性两个角度,ProxiMix均表现出显著有效性。
原文摘要 · Abstract (English)
Many bias mitigation methods have been developed for addressing fairness issues in machine learning. We found that using linear mixup alone, a data augmentation technique, for bias mitigation, can still retain biases present in dataset labels. Research presented in this paper aims to address this issue by proposing a novel pre-processing strategy in which both an existing mixup method and our new bias mitigation algorithm can be utilized to improve the generation of labels of augmented samples, which are proximity aware. Specifically, we proposed ProxiMix which keeps both pairwise and proximity relationships for fairer data augmentation. We conducted thorough experiments with three datasets, three ML models, and different hyperparameters settings. Our experimental results showed the effectiveness of ProxiMix from both fairness of predictions and fairness of recourse perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。