用合成人脸数据替代真实数据评估人脸识别,实现全流程隐私保护。
Benchmarking Face Recognition without Real Faces

- 构建合成数据集并对比真实基准,验证其评估有效性。
- MorphFace和Vec2Face表现最优,与真实基准一致性达自然差异范围。
- 适合关注隐私安全与模型公平评估的研究者参考。
合成人脸数据集已足以训练出性能媲美真实照片训练的面部识别模型,避免了收集真实生物特征数据带来的伦理与法律负担。然而,评估仍依赖真实人脸基准,使隐私问题仅解决一半。本文探讨合成数据能否替代真实基准进行评估。测试了12个合成数据集与7个成熟真实基准,使用24个涵盖卷积与变换器架构的预训练模型,评估内容包括生物特征验证指标、相似度分布、跨模型排名一致性及各数据集的分布特性。合成数据评估可靠性差异显著,其中表现最佳的MorphFace与Vec2Face能复现真实基准的相对行为,达成的一致性水平处于真实基准间原有分歧范围内。结果表明,精心设计的合成数据集可支持可靠的比较评估,推动面部识别训练与评估全流程向完全合成化与隐私保护迈进。
原文摘要 · Abstract (English)
Synthetic face datasets have become effective enough to train face recognition models with accuracy rivaling that of models trained on real photographs. This progress sidesteps the ethical and legal burdens of collecting real biometric data, yet evaluation has not kept pace. Even studies that train entirely on synthetic images still rely on real-face benchmarks to measure performance, leaving the privacy problem only half solved. We ask whether synthetic datasets can replace real benchmarks for face recognition evaluation. We test 12 synthetic datasets against 7 established real benchmarks using 24 pre-trained models that span both convolutional and transformer architectures. Our evaluation covers biometric verification metrics, similarity score distributions, cross-model ranking consistency, and the underlying distributional properties of each dataset. Benchmarking fidelity varies widely across the synthetic candidates, but the two strongest, MorphFace and Vec2Face, reproduce the relative behavior of real benchmarks and reach agreement levels that fall within the natural disagreement already observed among the real benchmarks themselves. These results establish that well-constructed synthetic datasets can support reliable comparative evaluation for face recognition, moving the field closer to a fully synthetic and privacy-preserving pipeline for both training and benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。