arXiv:2409.02867cs.CV2024-09ECCV被引 10

用扩散模型生成均衡数据能提升人脸识别准确率,但对公平性帮助有限。

The Impact of Balancing Real and Synthetic Data on Accuracy and Fairness in Face Recognition

  • 用扩散模型生成人口均衡的合成数据
  • 合成数据显著提升识别准确率,尤其与真实数据结合时
  • 平衡合成数据对公平性影响微弱,甚至可能恶化

近年来,深度人脸识别技术的发展推动了对大规模、多样化数据集的需求。然而,用于构建这些数据集的真实数据通常来自网络,往往因缺乏用户明确同意而引发严重隐私问题。此外,获取人口分布均衡的大规模数据集更加困难,因为不同群体图像的自然分布本身就存在不平衡。本文研究了人口均衡的真实数据与合成数据(单独或组合)对人脸识别模型准确率和公平性的影响。首先,采用多种生成方法平衡合成数据的人口代表性;随后,使用最先进的面部编码器在(合成与真实数据的组合)上训练并评估模型。研究发现:(i)基于扩散模型生成的数据在提升准确率方面效果显著,无论单独使用还是与真实数据子集结合;(ii)使用预训练生成模型提供的均衡数据对公平性几乎无改善作用,在几乎所有测试场景中,公平性得分保持不变或反而劣于非均衡真实数据集。源代码与数据已公开,支持可复现性。

原文摘要 · Abstract (English)

Over the recent years, the advancements in deep face recognition have fueled an increasing demand for large and diverse datasets. Nevertheless, the authentic data acquired to create those datasets is typically sourced from the web, which, in many cases, can lead to significant privacy issues due to the lack of explicit user consent. Furthermore, obtaining a demographically balanced, large dataset is even more difficult because of the natural imbalance in the distribution of images from different demographic groups. In this paper, we investigate the impact of demographically balanced authentic and synthetic data, both individually and in combination, on the accuracy and fairness of face recognition models. Initially, several generative methods were used to balance the demographic representations of the corresponding synthetic datasets. Then a state-of-the-art face encoder was trained and evaluated using (combinations of) synthetic and authentic images. Our findings emphasized two main points: (i) the increased effectiveness of training data generated by diffusion-based models in enhancing accuracy, whether used alone or combined with subsets of authentic data, and (ii) the minimal impact of incorporating balanced data from pre-trained generative methods on fairness (in nearly all tested scenarios using combined datasets, fairness scores remained either unchanged or worsened, even when compared to unbalanced authentic datasets). Source code and data are available at \url{https://cutt.ly/AeQy1K5G} for reproducibility.

人脸识别合成数据公平性扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。