arXiv:2507.11247cs.CVcs.LG2025-07被引 4

针对肤色等连续敏感属性,提出基于歧视程度的分组方法以发现隐蔽歧视。

Fairness-Aware Grouping for Continuous Sensitive Variables: Application for Debiasing Face Analysis with respect to Skin Tone

  • 按实际歧视水平分组,最大化组间歧视方差来识别关键子群体
  • 在CelebA和FFHQ数据集上发现比以往更细微的肤色歧视模式
  • 可直接用于模型后处理实现公平性提升,且不影响准确率

在法律框架下,公平性通常通过预定义分组计算差异影响或几率均等性等指标评估。然而,当敏感属性(如肤色)为连续变量时,固定分组可能忽略或掩盖少数子群体的歧视。为此,本文提出一种基于公平性的连续敏感属性分组方法,通过根据观测到的歧视水平进行分组,最大化基于组间歧视方差的新准则,从而识别最关键的子群体。我们在多个合成数据集上验证了该方法对人口分布变化的鲁棒性,揭示了歧视在敏感属性空间中的表现形式。此外,针对肤色场景引入单调公平性假设。在CelebA和FFHQ数据集上的实证结果表明,该分组方法能发现比以往更细致的歧视模式,且结果在不同数据集上对同一模型保持稳定。最后,我们利用该分组模型进行去偏,通过分组后处理预测公平分数。结果表明,该方法显著提升公平性,同时对准确率影响极小,验证了分组策略的有效性,并为工业应用铺平道路。

原文摘要 · Abstract (English)

Within a legal framework, fairness in datasets and models is typically assessed by dividing observations into predefined groups and then computing fairness measures (e.g., Disparate Impact or Equality of Odds with respect to gender). However, when sensitive attributes such as skin color are continuous, dividing into default groups may overlook or obscure the discrimination experienced by certain minority subpopulations. To address this limitation, we propose a fairness-based grouping approach for continuous (possibly multidimensional) sensitive attributes. By grouping data according to observed levels of discrimination, our method identifies the partition that maximizes a novel criterion based on inter-group variance in discrimination, thereby isolating the most critical subgroups. We validate the proposed approach using multiple synthetic datasets and demonstrate its robustness under changing population distributions - revealing how discrimination is manifested within the space of sensitive attributes. Furthermore, we examine a specialized setting of monotonic fairness for the case of skin color. Our empirical results on both CelebA and FFHQ, leveraging the skin tone as predicted by an industrial proprietary algorithm, show that the proposed segmentation uncovers more nuanced patterns of discrimination than previously reported, and that these findings remain stable across datasets for a given model. Finally, we leverage our grouping model for debiasing purpose, aiming at predicting fair scores with group-by-group post-processing. The results demonstrate that our approach improves fairness while having minimal impact on accuracy, thus confirming our partition method and opening the door for industrial deployment.

公平性肤色分析分组方法去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。