arXiv:2411.02188cs.CV2024-11中稿 · Publication in WAC…被引 9

用大模型提升合成人脸图像真实感,显著改善识别性能。

Digi2Real: Bridging the Realism Gap in Synthetic Data Face Recognition via Foundation Models

  • 结合图形管线与大模型,增强合成人脸图像的真实感。
  • 在多个基准上表现优于原始合成数据训练的模型。
  • 适合关注隐私保护与高效训练的人脸识别研究者。

近年来,面部识别系统的准确率大幅提升,主要得益于大规模数据集和神经网络架构的进步。然而,这些数据集往往未经明确同意收集,引发伦理与隐私问题。为此,有研究提出使用合成数据训练模型。但这类模型仍需真实数据训练生成模型,且性能通常不如真实数据训练的模型。例如,DigiFace 数据集通过图形管线生成身份和类内差异,不依赖真实数据训练模型,但在面部识别基准上表现不佳,可能因图像缺乏真实感。本文提出一种新颖的现实感迁移框架,利用大规模面部基础模型提升合成图像的真实感。通过将图形管线的可控性与我们的现实感增强技术相结合,生成大量逼真的变化。实证评估表明,使用增强后数据训练的模型在性能上显著优于基线。代码与数据集将公开提供。

原文摘要 · Abstract (English)

The accuracy of face recognition systems has improved significantly in the past few years, thanks to the large amount of data collected and advancements in neural network architectures. However, these large-scale datasets are often collected without explicit consent, raising ethical and privacy concerns. To address this, there have been proposals to use synthetic datasets for training face recognition models. Yet, such models still rely on real data to train the generative models and generally exhibit inferior performance compared to those trained on real datasets. One of these datasets, DigiFace, uses a graphics pipeline to generate different identities and intra-class variations without using real data in model training. However, the performance of this approach is poor on face recognition benchmarks, possibly due to the lack of realism in the images generated by the graphics pipeline. In this work, we introduce a novel framework for realism transfer aimed at enhancing the realism of synthetically generated face images. Our method leverages the large-scale face foundation model, and we adapt the pipeline for realism enhancement. By integrating the controllable aspects of the graphics pipeline with our realism enhancement technique, we generate a large amount of realistic variations, combining the advantages of both approaches. Our empirical evaluations demonstrate that models trained using our enhanced dataset significantly improve the performance of face recognition systems over the baseline. The source code and dataset will be publicly accessible at the following link: https://www.idiap.ch/paper/digi2real

人脸识别合成数据基础模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。