arXiv:2507.20782cs.CVcs.AI2025-07中稿 · publication in IEE…被引 3

用合成数据训练人脸识别模型,能兼顾准确率与公平性。

Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data

  • 构建平衡数据集FairFaceGen,结合文本生成图像技术
  • 合成数据在部分基准上接近真实数据,且更易降低种族偏见
  • 高质量的类别内增广能显著提升识别精度与公平性

合成数据已成为训练人脸识别(FR)模型的有前景替代方案,具备可扩展性、隐私合规性及潜在的偏见缓解优势。然而,是否能在合成数据上同时实现高准确率与公平性仍存疑问。本文通过两种先进文本到图像生成器Flux.1-dev和Stable Diffusion v3.5(SD35),生成平衡人脸数据集FairFaceGen,并结合多种身份增广方法(Arc2Face及四种IP-Adapters)。为确保公平比较,合成与真实数据集保持相等身份数量。在标准(LFW、AgeDB-30等)和挑战性基准(IJB-B/C)上评估性能,在RFW数据集上检测偏见。结果表明:尽管合成数据在IJB-B/C上的泛化能力仍落后于真实数据,但经过均衡设计的合成数据集,尤其是由SD35生成的,展现出偏见缓解潜力。此外,类内增广的数量与质量显著影响识别准确率与公平性。这些发现为利用合成数据构建更公平的FR系统提供了实用指导。

原文摘要 · Abstract (English)

Synthetic data has emerged as a promising alternative for training face recognition (FR) models, offering advantages in scalability, privacy compliance, and potential for bias mitigation. However, critical questions remain on whether both high accuracy and fairness can be achieved with synthetic data. In this work, we evaluate the impact of synthetic data on bias and performance of FR systems. We generate balanced face dataset, FairFaceGen, using two state of the art text-to-image generators, Flux.1-dev and Stable Diffusion v3.5 (SD35), and combine them with several identity augmentation methods, including Arc2Face and four IP-Adapters. By maintaining equal identity count across synthetic and real datasets, we ensure fair comparisons when evaluating FR performance on standard (LFW, AgeDB-30, etc.) and challenging IJB-B/C benchmarks and FR bias on Racial Faces in-the-Wild (RFW) dataset. Our results demonstrate that although synthetic data still lags behind the real datasets in the generalization on IJB-B/C, demographically balanced synthetic datasets, especially those generated with SD35, show potential for bias mitigation. We also observe that the number and quality of intra-class augmentations significantly affect FR accuracy and fairness. These findings provide practical guidelines for constructing fairer FR systems using synthetic data.

人脸识别合成数据偏见缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。