提出公平性导向的标签噪声修正方法,提升数据集的群体平等性。
Fair-OBNC: Correcting Label Noise for Fairer Datasets
- 基于集成模型误差与群体公平性提升潜力,动态调整标签修正顺序。
- 修正后数据训练的模型,群体公平性平均提升150%。
- 适合关注算法公平性的研究者与实际应用开发者。
自动化决策系统使用的数据常反映历史歧视行为,此类偏差常与标签噪声相关,如COMPAS数据集中非洲裔被告被错误标记为高再犯风险的比例高于白人。在这些有偏数据上训练的模型可能延续甚至加剧性别、种族或年龄等敏感信息上的偏见。尽管已有多种标签噪声修正方法,但均仅关注模型性能。本文提出Fair-OBNC,一种考虑公平性的标签噪声修正方法,旨在生成具有可测量群体平等性的训练数据。该方法改进了基于排序的噪声修正(OBNC),通过结合集成模型的误差边际与数据集潜在的群体公平性提升,动态调整标签修正优先级。在不同控制标签噪声场景下,与多种预处理方法对比,结果表明:所提方法在整体上优于现有方法,能更准确还原原始标签;使用修正数据训练的模型,其群体公平性平均较含噪声数据训练模型提升150%。
原文摘要 · Abstract (English)
Data used by automated decision-making systems, such as Machine Learning models, often reflects discriminatory behavior that occurred in the past. These biases in the training data are sometimes related to label noise, such as in COMPAS, where more African-American offenders are wrongly labeled as having a higher risk of recidivism when compared to their White counterparts. Models trained on such biased data may perpetuate or even aggravate the biases with respect to sensitive information, such as gender, race, or age. However, while multiple label noise correction approaches are available in the literature, these focus on model performance exclusively. In this work, we propose Fair-OBNC, a label noise correction method with fairness considerations, to produce training datasets with measurable demographic parity. The presented method adapts Ordering-Based Noise Correction, with an adjusted criterion of ordering, based both on the margin of error of an ensemble, and the potential increase in the observed demographic parity of the dataset. We evaluate Fair-OBNC against other different pre-processing techniques, under different scenarios of controlled label noise. Our results show that the proposed method is the overall better alternative within the pool of label correction methods, being capable of attaining better reconstructions of the original labels. Models trained in the corrected data have an increase, on average, of 150% in demographic parity, when compared to models trained in data with noisy labels, across the considered levels of label noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。