arXiv:2511.17393cs.CVcs.AI2025-11

生成多样化公平的面部数据集,提升人脸识别公正性

Designing and Generating Diverse, Equitable Face Image Datasets for Face Verification Tasks

  • 用生成模型合成高质多样人脸图像,覆盖多族裔性别特征
  • 构建含926人、27780张图的DIF-V数据集,用于验证模型公平性
  • 发现现有模型对特定性别种族有偏见,风格修改会降低性能

人脸验证在在线银行和设备安全访问等场景中至关重要。现有面部数据集普遍存在种族、性别等人口统计学偏差,限制了验证系统的有效性与公平性。为此,我们提出一种整合先进生成模型的综合方法,生成高质量、多样化的合成人脸图像,强调涵盖多种面部特征,并符合身份证照片的合规要求。同时,我们发布了面向验证任务的多样包容人脸数据集DIF-V,包含926个唯一身份的27,780张图像,作为未来研究基准。分析显示,现有验证模型对特定性别和种族存在偏差,且应用身份风格修改会显著降低模型性能。本工作不仅推动人工智能中多样性与伦理的讨论,也为开发更包容可靠的面部验证技术奠定基础。

原文摘要 · Abstract (English)

Face verification is a significant component of identity authentication in various applications including online banking and secure access to personal devices. The majority of the existing face image datasets often suffer from notable biases related to race, gender, and other demographic characteristics, limiting the effectiveness and fairness of face verification systems. In response to these challenges, we propose a comprehensive methodology that integrates advanced generative models to create varied and diverse high-quality synthetic face images. This methodology emphasizes the representation of a diverse range of facial traits, ensuring adherence to characteristics permissible in identity card photographs. Furthermore, we introduce the Diverse and Inclusive Faces for Verification (DIF-V) dataset, comprising 27,780 images of 926 unique identities, designed as a benchmark for future research in face verification. Our analysis reveals that existing verification models exhibit biases toward certain genders and races, and notably, applying identity style modifications negatively impacts model performance. By tackling the inherent inequities in existing datasets, this work not only enriches the discussion on diversity and ethics in artificial intelligence but also lays the foundation for developing more inclusive and reliable face verification technologies

人脸验证数据集公平性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。