提出面向下游任务的公平预处理方法,兼顾公平性与模型性能。
Task-tailored Pre-processing: Fair Downstream Supervised Learning
- 设计任务定制化预处理流程,平衡公平性与预测效用。
- 理论证明变换后数据可提升多种下游模型的公平性且不损失性能。
- 在表格和图像数据上验证,尤其对视觉任务仅调整必要语义特征。
针对监督学习中的算法公平性问题,本文探讨了预处理阶段的公平性方法。现有方法分为数据公平性和任务定制公平性两类:前者忽略下游模型类型,强制各群体间预测一致;后者则考虑具体学习任务。本文指出数据公平性方法在HGR相关性视角下正则化过强,因此提出一种新的任务定制化预处理方法,显式权衡公平性与效用。通过理论分析,给出任意下游监督模型在变换数据上实现公平性提升且保持性能的充分条件。据我们所知,这是首个对任务定制方法进行下游保障理论分析的工作。在表格和图像数据集上的对比实验表明,该框架在多种下游模型中均能保持稳定公平-效用折衷,尤其在计算机视觉任务中仅调整与核心任务相关的必要语义特征即可达成公平。
原文摘要 · Abstract (English)
Fairness-aware machine learning has recently attracted various communities to mitigate discrimination against certain societal groups in data-driven tasks. For fair supervised learning, particularly in pre-processing, there have been two main categories: data fairness and task-tailored fairness. The former directly finds an intermediate distribution among the groups, independent of the type of the downstream model, so a learned downstream classification/regression model returns similar predictive scores to individuals inputting the same covariates irrespective of their sensitive attributes. The latter explicitly takes the supervised learning task into account when constructing the pre-processing map. In this work, we study algorithmic fairness for supervised learning and argue that the data fairness approaches impose overly strong regularization from the perspective of the HGR correlation. This motivates us to devise a novel pre-processing approach tailored to supervised learning. We account for the trade-off between fairness and utility in obtaining the pre-processing map. Then we study the behavior of arbitrary downstream supervised models learned on the transformed data to find sufficient conditions to guarantee their fairness improvement and utility preservation. To our knowledge, no prior work in the branch of task-tailored methods has theoretically investigated downstream guarantees when using pre-processed data. We further evaluate our framework through comparison studies based on tabular and image data sets, showing the superiority of our framework which preserves consistent trade-offs among multiple downstream models compared to recent competing models. Particularly for computer vision data, we see our method alters only necessary semantic features related to the central machine learning task to achieve fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。