arXiv:2409.01928cs.CVcs.AI2024-09中稿 · the 27th Internati…被引 3

提出新指标评估人脸识别模型的群体偏见,更全面反映真实场景中的公平性问题。

Comprehensive Equity Index (CEI): Definition and Application to Bias Evaluation in Biometrics

  • 融合分布整体形状与尾部概率,综合评估模型偏差
  • 在NIST FRVT测试中验证,能有效识别真实数据中的偏见
  • 适合用于高精度人脸识别系统的公平性评估

我们提出一种新型度量方法,旨在量化机器学习模型的偏见行为。该度量的核心是一种新的分数分布相似性度量,兼顾分布的整体形状和尾部概率。在实际应用中,我们聚焦于人脸识别系统的运行评估,特别关注人口统计学偏见的量化,此应用场景下该度量尤为适用。近年来,生物识别系统中的人口统计学偏见与公平性问题备受关注。随着这些系统在社会中的广泛应用,不同群体受到的对待是否公平引发担忧。预防和缓解偏见的关键一步是首先检测并量化其存在。传统上,有两种方法被用于衡量群体间的差异:1)误差率差异;2)识别得分分布差异。本文提出的综合公平指数(CEI)在两者间进行权衡,同时考虑分布尾部的误差和整体形状。该度量在真实世界场景中表现良好,基于包含多种协变量和人口群体的高精度系统与真实人脸数据库的NIST FRVT评估。我们首先揭示了现有度量在现实设置中正确评估偏见能力的局限性,随后提出了新度量以克服这些缺陷。我们在两个最先进的模型和四个广泛使用的数据库上测试该度量,证明其能够克服以往偏见度量的主要缺陷。

原文摘要 · Abstract (English)

We present a novel metric designed, among other applications, to quantify biased behaviors of machine learning models. As its core, the metric consists of a new similarity metric between score distributions that balances both their general shapes and tails' probabilities. In that sense, our proposed metric may be useful in many application areas. Here we focus on and apply it to the operational evaluation of face recognition systems, with special attention to quantifying demographic biases; an application where our metric is especially useful. The topic of demographic bias and fairness in biometric recognition systems has gained major attention in recent years. The usage of these systems has spread in society, raising concerns about the extent to which these systems treat different population groups. A relevant step to prevent and mitigate demographic biases is first to detect and quantify them. Traditionally, two approaches have been studied to quantify differences between population groups in machine learning literature: 1) measuring differences in error rates, and 2) measuring differences in recognition score distributions. Our proposed Comprehensive Equity Index (CEI) trade-offs both approaches combining both errors from distribution tails and general distribution shapes. This new metric is well suited to real-world scenarios, as measured on NIST FRVT evaluations, involving high-performance systems and realistic face databases including a wide range of covariates and demographic groups. We first show the limitations of existing metrics to correctly assess the presence of biases in realistic setups and then propose our new metric to tackle these limitations. We tested the proposed metric with two state-of-the-art models and four widely used databases, showing its capacity to overcome the main flaws of previous bias metrics.

人脸识别公平性评估偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。