用合成数据校准人脸识别阈值,发现其在真实场景下效果不稳定。
On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems

- 用合成人脸数据模拟真实验证场景,测试阈值校准效果
- 合成数据在低误匹配率下表现差,尾部得分分布不匹配
- 高安全场景需真实数据验证,合成数据仅作初步参考
新部署的入境/出境系统(EES)将大规模生物特征验证引入欧洲边境管控,要求人脸识别系统在极低误匹配率(FMR)下运行。尽管监管框架设定了中央系统的性能目标,但未明确成员国层面如何实际校准验证阈值。在实际操作中,因法律、物流和隐私限制,获取代表性真实数据用于校准常受制约。本文研究了合成人脸数据在边境管控中证件对活体验证场景下的阈值校准应用。分析了合成与真实数据在真用户与伪造者得分分布上的对齐情况,并评估了校准阈值在不同域间的迁移能力,重点关注低FMR工作点。结果表明,合成数据在受控环境下可近似校准行为,但在非受限条件下无法可靠泛化,因得分分布尾部存在偏差,导致识别性能显著下降且更易遭受变脸攻击。我们进一步证明,即使在合成数据间,校准结果也高度依赖数据集。总体而言,虽然合成数据对系统开发和初步校准有帮助,但高安全性部署中的可靠阈值选择仍需使用具有代表性的真实数据进行验证与调整。
原文摘要 · Abstract (English)
The recently deployed Entry/Exit System (EES) introduces large-scale biometric verification into European border control, requiring face recognition systems to operate at extremely low false match rates (FMR). While regulatory frameworks define performance targets at the EES Central System level, they do not specify how verification thresholds should be calibrated in practice at the Member State level. In operational settings, obtaining representative real-world data for calibration is often constrained by legal, logistical, and privacy limitations. In this work, we investigate the use of synthetic face data for threshold calibration in document-to-live verification scenarios relevant to border control systems. We analyze the alignment of genuine and impostor score distributions between synthetic and real datasets and evaluate the transferability of calibrated thresholds across domains, with a focus on low-FMR operating points. Our results show that synthetic data can approximate calibration behavior in controlled settings, but fails to reliably generalize to unconstrained conditions due to mismatches in score distribution tails. These discrepancies lead to significant degradation in recognition performance and increased vulnerability to morph-based attacks. We further demonstrate that calibration outcomes are highly dataset-dependent, even across synthetic datasets. Overall, our findings highlight that while synthetic data is useful for system development and preliminary calibration, our results indicate that reliable threshold selection in high-security deployments typically requires validation and adjustment using representative real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。