高准确率骗人,人脸识别在不同人群表现差异大
Why Aggregate Accuracy is Inadequate for Evaluating Fairness in Law Enforcement Facial Recognition Systems
- 用子群体错误率揭示整体准确率掩盖的不公平性
- 相同整体准确率下,不同人群误判率差异显著
- 适合关注算法公平性与执法风险的研究者
人脸识别系统在执法和安保领域日益普及,其算法决策可能带来重大社会影响。尽管报告的总体准确率较高,但越来越多证据表明,这些系统在不同人口群体中表现不均,导致误判率差异显著,可能造成严重后果。本文指出,仅依赖整体准确率无法有效评估人脸识别系统的公平性和可靠性。通过分析子群体层面的错误分布(包括假阳性率FPR和假阴性率FNR),发现整体性能指标会掩盖关键的群体差异。实证显示,具有相似整体准确率的系统,其子群体错误率可能截然不同。论文进一步探讨了以准确率为导向的评估方式在执法应用中的操作风险,如错误识别可能导致无辜怀疑或遗漏识别。强调需采用注重公平性的评估方法和模型无关的审计策略,实现对真实部署系统的事后评估。研究呼吁摒弃单一准确率指标,建立更全面的评估框架以推动负责任的人工智能部署。
原文摘要 · Abstract (English)
Facial recognition systems are increasingly deployed in law enforcement and security contexts, where algorithmic decisions can carry significant societal consequences. Despite high reported accuracy, growing evidence demonstrates that such systems often exhibit uneven performance across demographic groups, leading to disproportionate error rates and potential harm. This paper argues that aggregate accuracy is an insufficient metric for evaluating the fairness and reliability of facial recognition systems in high-stakes environments. Through analysis of subgroup-level error distribution, including false positive rate (FPR) and false negative rate (FNR), the paper demonstrates how aggregate performance metrics can obscure critical disparities across demographic groups. Empirical observations show that systems with similar overall accuracy can exhibit substantially different fairness profiles, with subgroup error rates varying significantly despite a single aggregate metric. The paper further examines the operational risks associated with accuracy-centric evaluation practices in law enforcement applications, where misclassification may result in wrongful suspicion or missed identification. It highlights the importance of fairness-aware evaluation approaches and model-agnostic auditing strategies that enable post-deployment assessment of real-world systems. The findings emphasise the need to move beyond accuracy as a primary metric and adopt more comprehensive evaluation frameworks for responsible AI deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。