用最优传输修复数据偏见,让少数群体数据更公平。
Overcoming Representation Bias in Fairness-Aware data Repair using Optimal Transport
- 基于贝叶斯非参数停顿规则学习各属性子群分布,缓解代表性不足问题。
- 修复后数据在模拟与基准数据集上表现优异,公平性提升且损失可控。
- 可应用于未参与训练的归档数据,适合需要公平性保障的实际场景。
最优传输(OT)在以公平方式转换数据分布方面具有重要作用。通常,OT 算子从带有不公平属性标签的数据中学习,用于数据修复。但该方法存在两大局限:(i) 对代表性不足子群的 OT 算子学习效果差(易受表示偏见影响);(ii) 无法对同分布但未参与训练的归档数据进行修复。本文通过采用贝叶斯非参数停顿规则学习每个属性标签数据分布的组件,解决上述问题。由此得到的最优传输量化算子可用于修复归档数据。我们提出了一种新的公平分布目标定义,并引入可量化指标,实现公平性与数据失真之间的权衡。实验表明,所提方法在模拟与基准数据集上均表现出色,具备较强的抗表示偏见能力。
原文摘要 · Abstract (English)
Optimal transport (OT) has an important role in transforming data distributions in a manner which engenders fairness. Typically, the OT operators are learnt from the unfair attribute-labelled data, and then used for their repair. Two significant limitations of this approach are as follows: (i) the OT operators for underrepresented subgroups are poorly learnt (i.e. they are susceptible to representation bias); and (ii) these OT repairs cannot be effected on identically distributed but out-of-sample (i.e.\ archival) data. In this paper, we address both of these problems by adopting a Bayesian nonparametric stopping rule for learning each attribute-labelled component of the data distribution. The induced OT-optimal quantization operators can then be used to repair the archival data. We formulate a novel definition of the fair distributional target, along with quantifiers that allow us to trade fairness against damage in the transformed data. These are used to reveal excellent performance of our representation-bias-tolerant scheme in simulated and benchmark data sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。