arXiv:2506.01532cs.CV2025-06

用连续标签重新定义种族,让人脸识别更公平

Balancing Beyond Discrete Categories: Continuous Demographic Labels for Fair Face Recognition

  • 将种族标签从离散类别改为连续变量,更真实反映数据分布
  • 连续空间平衡的数据集训练出的模型性能显著优于离散平衡
  • 适合关注公平性、数据偏见研究的研究者参考

面部识别模型中的偏见问题长期存在。以往研究多从模型或数据角度出发,但对数据偏见的缓解方式有限,且未能深入理解问题本质。本文提出将种族标签视为连续变量而非每个身份的离散类别。通过实验与理论验证,我们发现同一种族内的不同身份对数据平衡的贡献不均,因此相同数量的身份并不意味着数据平衡。在连续空间中平衡数据集后,训练的模型性能持续优于在离散空间中平衡的数据集。我们共训练了65个以上模型,并构建了20多个原始数据集的子集。

原文摘要 · Abstract (English)

Bias has been a constant in face recognition models. Over the years, researchers have looked at it from both the model and the data point of view. However, their approach to mitigation of data bias was limited and lacked insight on the real nature of the problem. Here, in this document, we propose to revise our use of ethnicity labels as a continuous variable instead of a discrete value per identity. We validate our formulation both experimentally and theoretically, showcasing that not all identities from one ethnicity contribute equally to the balance of the dataset; thus, having the same number of identities per ethnicity does not represent a balanced dataset. We further show that models trained on datasets balanced in the continuous space consistently outperform models trained on data balanced in the discrete space. We trained more than 65 different models, and created more than 20 subsets of the original datasets.

人脸识别公平性数据偏见连续标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。