arXiv:2509.08163cs.LGq-fin.RM2025-09被引 2

通过正则化减少模型对多类型敏感属性的依赖,提升公平性。

Machine Learning with Multitype Protected Attributes: Intersectional Fairness through Regularisation

  • 用距离协方差正则化降低预测与敏感属性关联
  • 在回归和分类任务中同时处理多属性公平性问题
  • 可应对交叉群体差异,适合保险、招聘等场景

机器学习中的公平性问题在性别、种族等保护属性上尤为关键。现有研究多聚焦于二分类,但回归任务(如保险定价、招聘评分)同样需保障公平性。反歧视法律也涵盖连续属性(如年龄),而多数方法不适用。现实中多个保护属性常共存,但现有方法常忽视‘公平性分选’问题,即忽略交叉子群体(如非裔女性或拉丁裔男性)间的差异。本文提出基于距离协方差正则化的框架,缓解模型预测与保护属性之间的关联,符合人口均等性定义,并能捕捉线性和非线性依赖。为适应多属性场景,引入联合距离协方差(JdCov)与新提出的拼接距离协方差(CCdCov),有效解决回归与分类任务中的交叉公平性问题。我们还讨论了正则化强度的校准方法,包括基于熵散度(Jensen-Shannon divergence)的分布差异量化。实验在COMPAS再犯数据集和大型汽车保险索赔数据集上验证了有效性。

原文摘要 · Abstract (English)

Ensuring equitable treatment (fairness) across protected attributes (such as gender or ethnicity) is a critical issue in machine learning. Most existing literature focuses on binary classification, but achieving fairness in regression tasks-such as insurance pricing or hiring score assessments-is equally important. Moreover, anti-discrimination laws also apply to continuous attributes, such as age, for which many existing methods are not applicable. In practice, multiple protected attributes can exist simultaneously; however, methods targeting fairness across several attributes often overlook so-called "fairness gerrymandering", thereby ignoring disparities among intersectional subgroups (e.g., African-American women or Hispanic men). In this paper, we propose a distance covariance regularisation framework that mitigates the association between model predictions and protected attributes, in line with the fairness definition of demographic parity, and that captures both linear and nonlinear dependencies. To enhance applicability in the presence of multiple protected attributes, we extend our framework by incorporating two multivariate dependence measures based on distance covariance: the previously proposed joint distance covariance (JdCov) and our novel concatenated distance covariance (CCdCov), which effectively address fairness gerrymandering in both regression and classification tasks involving protected attributes of various types. We discuss and illustrate how to calibrate regularisation strength, including a method based on Jensen-Shannon divergence, which quantifies dissimilarities in prediction distributions across groups. We apply our framework to the COMPAS recidivism dataset and a large motor insurance claims dataset.

公平性正则化多属性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。