提出基于密度估计的重加权方法,提升回归模型的公平性分离性。
FairReweighing: Density Estimation-Based Reweighing Framework for Improving Separation in Fair Regression
- 通过密度估计重构数据权重,确保满足公平分离准则
- 在真实与合成数据上显著优于现有方法,同时保持高准确率
- 适用于需要公平性的回归任务,尤其关注种族性别等敏感属性
AI 软件在高风险公共与工业场景中广泛应用,但其缺乏透明性引发对不同种族、性别、年龄群体是否公平的担忧。尽管已有大量公平性研究,但多集中于二分类任务,回归任务中的公平性仍被忽视。本文采用基于互信息的度量来评估分离性违规,并扩展该度量以适配分类与回归任务,支持二值与连续敏感属性。受公平分类中重加权算法启发,提出 FairReweighing 预处理方法,基于密度估计调整样本权重,使学习模型满足分离性标准。理论证明在数据独立假设下可保证训练数据满足分离性。实验在合成与真实数据集上显示,FairReweighing 在提升分离性的同时保持高精度,优于当前最先进的回归公平性方法。
原文摘要 · Abstract (English)
There has been a prevalence of applying AI software in both high-stakes public-sector and industrial contexts. However, the lack of transparency has raised concerns about whether these data-informed AI software decisions secure fairness against people of all racial, gender, or age groups. Despite extensive research on emerging fairness-aware AI software, up to now most efforts to solve this issue have been dedicated to binary classification tasks. Fairness in regression is relatively underexplored. In this work, we adopted a mutual information-based metric to assess separation violations. The metric is also extended so that it can be directly applied to both classification and regression problems with both binary and continuous sensitive attributes. Inspired by the Reweighing algorithm in fair classification, we proposed a FairReweighing pre-processing algorithm based on density estimation to ensure that the learned model satisfies the separation criterion. Theoretically, we show that the proposed FairReweighing algorithm can guarantee separation in the training data under a data independence assumption. Empirically, on both synthetic and real-world data, we show that FairReweighing outperforms existing state-of-the-art regression fairness solutions in terms of improving separation while maintaining high accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。