不同公平性指标评估结果常冲突,单一指标不可靠
When Fairness Metrics Disagree: Evaluating the Reliability of Demographic Fairness Assessment in Machine Learning
- 用多指标分析人脸识别中的群体偏差,对比多种公平性度量
- 同一模型在不同指标下评估结果矛盾,结论可能完全相反
- 提出公平性分歧指数FDI,适合关注评估可靠性的研究者
机器学习系统的公平性评估已成为生物特征识别、医疗决策和自动风险评估等高风险应用的核心议题。现有方法通常依赖少数公平性指标来评估模型在不同群体间的表现,隐含假设这些指标能给出一致可靠的结论。然而,不同公平性指标反映模型性能的不同统计特性,可能对同一系统产生相互冲突的评估结果。本文通过系统性的多指标分析,研究了机器学习模型中人口群体偏差的评估一致性。以人脸识别为受控实验场景,在多种群体划分下,使用误差率差异和基于性能的度量等常用公平性指标评估模型表现。结果表明,公平性评估结果随指标选择显著变化,导致关于模型偏差的结论相互矛盾。为此,我们提出公平性分歧指数(FDI),用于量化不同指标间的一致性程度。进一步发现,分歧在不同阈值和模型配置下依然较高。这些结果揭示了当前公平性评估实践的关键局限,表明仅依赖单一指标无法实现可靠的偏差评估。
原文摘要 · Abstract (English)
The evaluation of fairness in machine learning systems has become a central concern in high-stakes applications, including biometric recognition, healthcare decision-making, and automated risk assessment. Existing approaches typically rely on a small number of fairness metrics to assess model behaviour across group partitions, implicitly assuming that these metrics provide consistent and reliable conclusions. However, different fairness metrics capture distinct statistical properties of model performance and may therefore produce conflicting assessments when applied to the same system. In this work, we investigate the consistency of fairness evaluation by conducting a systematic multi-metric analysis of demographic bias in machine learning models. Using face recognition as a controlled experimental setting, we evaluate model performance across multiple group partitions under a range of commonly used fairness metrics, including error-rate disparities and performance-based measures. Our results demonstrate that fairness assessments can vary significantly depending on the choice of metrics, leading to contradictory conclusions regarding model bias. To quantify this phenomenon, we introduce the Fairness Disagreement Index (FDI), a measure designed to capture the degree of inconsistency across fairness metrics. We further show that disagreement remains high across thresholds and model configurations. These findings highlight a critical limitation in current fairness evaluation practices and suggest that single-metric reporting is insufficient for reliable bias assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。