用自监督学习构建通用人脸安全模型,提升伪造检测泛化能力。
FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation Learning

- 通过掩码图像建模与实例判别结合,学习人脸局部与全局特征。
- 在10个数据集上超越监督预训练和现有自监督方法,跨域检测性能领先。
- 适合需要高泛化能力的面部安全任务,如深度伪造、活体检测等。
本文探讨如何利用大量未标注真实人脸数据,学习鲁棒且可迁移的人脸表征以提升各类人脸安全任务的泛化性能。提出首个自监督预训练框架FSFM,融合掩码图像建模(MIM)与实例判别(ID),设计简单的CRFR-P掩码策略,强制模型捕捉区域内部一致性与跨区域连贯性。同时构建与MIM协同的ID网络,通过自蒸馏建立局部到全局对应关系。三个学习目标(3C)共同编码真实人脸的局部特征与全局语义。预训练后,采用基础ViT作为通用视觉基础模型,应用于跨数据集深度伪造检测、跨域人脸反欺骗及未知扩散伪造检测。在10个公开数据集上的实验证明,该模型在迁移性能上优于监督预训练、视觉与人脸自监督学习方法,甚至超越特定任务的最新技术。
原文摘要 · Abstract (English)
This work asks: with abundant, unlabeled real faces, how to learn a robust and transferable facial representation that boosts various face security tasks with respect to generalization performance? We make the first attempt and propose a self-supervised pretraining framework to learn fundamental representations of real face images, FSFM, that leverages the synergy between masked image modeling (MIM) and instance discrimination (ID). We explore various facial masking strategies for MIM and present a simple yet powerful CRFR-P masking, which explicitly forces the model to capture meaningful intra-region consistency and challenging inter-region coherency. Furthermore, we devise the ID network that naturally couples with MIM to establish underlying local-to-global correspondence via tailored self-distillation. These three learning objectives, namely 3C, empower encoding both local features and global semantics of real faces. After pretraining, a vanilla ViT serves as a universal vision foundation model for downstream face security tasks: cross-dataset deepfake detection, cross-domain face anti-spoofing, and unseen diffusion facial forgery detection. Extensive experiments on 10 public datasets demonstrate that our model transfers better than supervised pretraining, visual and facial self-supervised learning arts, and even outperforms task-specialized SOTA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。