arXiv:2505.00830cs.LG2025-05被引 2

提出新公平性度量,评估回归模型在多重身份组合下的偏差。

Intersectional Divergence: Measuring Fairness in Regression

  • 通过交集偏差(ID)衡量多个保护属性组合下的回归公平性。
  • 考虑用户关注的目标范围,区分预测在关键区间的偏差影响。
  • 可转为损失函数IDLoss,提升公平性且不牺牲预测性能。

机器学习公平性研究多聚焦分类任务,忽视了回归任务的关键缺口。本文提出一种新型回归公平性度量方法——交集偏差(Intersectional Divergence, ID),突破以往仅关注单一保护属性的局限,全面考量所有保护属性的组合影响。同时指出,仅评估群体平均误差不足以反映真实偏差,尤其当目标域存在分布偏倚时。为此,ID不仅描述多属性交叉下模型行为的公平性,还强调预测在用户最关心的目标区间内的影响。我们进一步将ID扩展为可优化的损失函数IDLoss,具备收敛保证与分段光滑特性,便于实际训练。大量实验表明,ID能揭示模型行为与公平性的独特洞察;引入IDLoss可显著提升单属性及交集层面的公平性,同时保持良好的预测表现。

原文摘要 · Abstract (English)

Fairness in machine learning research is commonly framed in the context of classification tasks, leaving critical gaps in regression. In this paper, we propose a novel approach to measure intersectional fairness in regression tasks, going beyond the focus on single protected attributes from existing work to consider combinations of all protected attributes. Furthermore, we contend that it is insufficient to measure the average error of groups without regard for imbalanced domain preferences. Accordingly, we propose Intersectional Divergence (ID) as the first fairness measure for regression tasks that 1) describes fair model behavior across multiple protected attributes and 2) differentiates the impact of predictions in target ranges most relevant to users. We extend our proposal demonstrating how ID can be adapted into a loss function, IDLoss, that satisfies convergence guarantees and has piecewise smooth properties that enable practical optimization. Through an extensive experimental evaluation, we demonstrate how ID allows unique insights into model behavior and fairness, and how incorporating IDLoss into optimization can considerably improve single-attribute and intersectional model fairness while maintaining a competitive balance in predictive performance.

回归公平性交集偏差模型优化数据偏倚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。