arXiv:2606.31704cs.CVcs.LG2026-06

标注人脸性别与族裔,评估检测模型公平性

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

论文配图:WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation
图 1 · 摘自论文原文
  • 在WIDER-FACE基础上人工标注人脸族裔与性别
  • 黑人面孔检测准确率显著偏低,剔除该群体训练加剧不公平
  • 适合研究人脸识别公平性与偏见缓解的学者使用

面部检测模型在真实应用中存在跨人群性能差异的公平性问题。由于缺乏带敏感特征标注的数据集,相关研究受限。为此,我们推出WIDER-FAIR,基于广泛使用的WIDER-FACE基准,手动标注了16,256张图像中每张人脸的感知族裔(亚裔、黑人、印度人、白人)和性别。通过人脸嵌入、K近邻分类器和t-SNE可视化验证了标注一致性。以YOLOv5为例进行消融实验,发现黑人面孔检测性能明显更低,且从训练集中移除黑人群体将比其他族裔更严重加剧公平性偏差。结果表明,带人口统计学标注的数据集对理解与评估面部检测模型中的偏见具有重要价值。

原文摘要 · Abstract (English)

The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance disparities across demographic groups. A key obstacle to studying and mitigating such biases is the lack of face detection datasets with sensitive feature annotations. To address this gap, we introduce WIDER-FAIR, a new dataset built on the widely used WIDER-FACE benchmark, manually annotated with the perceived ethnicity and sex of each face. The dataset contains 16,256 images annotated across four ethnic groups: Asian, Black, Indian, and White, and two sex categories. We assess the quality and coherence of the annotations using face embeddings, a K-Nearest Neighbors classifier, and a t-SNE visualization, all of which support the consistency of the labeling process. As a demonstration of the dataset's potential, we train a YOLOv5 model and perform ablation studies on each sensitive feature. Among other findings, our experiments show that detection performance is notably lower for faces of Black individuals, and that excluding this group from training increases fairness disparity more than excluding any other ethnic group. These observations illustrate the value of demographically annotated datasets for understanding and evaluating bias in face detection models.

人脸检测公平性数据标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。