用图文模型让人脸换脸攻击检测可解释,一眼看懂为何可疑。
Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model
- 基于CLIP的零样本框架,自动描述攻击特征。
- 在5种生成方式、3种介质上检测准确率超主流模型。
- 适合需要可解释性的人脸验证系统开发者。
人脸换脸攻击检测已成为确保身份验证可靠性的重要环节。本文提出一种多模态学习方法,能为换脸攻击提供文本描述。我们首次证明,使用对比语言图像预训练(CLIP)进行零样本评估,不仅能实现泛化性强的换脸攻击检测,还能精准预测最相关的文本片段。针对十种不同长度的文本提示进行了深入分析,这些提示均基于人类可理解的描述设计。实验基于公开人脸生物特征数据集构建的换脸数据集展开,评估了当前最优预训练神经网络与所提框架在五种不同换脸生成技术、三种不同介质下的零样本表现。
原文摘要 · Abstract (English)
Morphing attack detection has become an essential component of face recognition systems for ensuring a reliable verification scenario. In this paper, we present a multimodal learning approach that can provide a textual description of morphing attack detection. We first show that zero-shot evaluation of the proposed framework using Contrastive Language-Image Pretraining (CLIP) can yield not only generalizable morphing attack detection, but also predict the most relevant text snippet. We present an extensive analysis of ten different textual prompts that include both short and long textual prompts. These prompts are engineered by considering the human understandable textual snippet. Extensive experiments were performed on a face morphing dataset that was developed using a publicly available face biometric dataset. We present an evaluation of SOTA pre-trained neural networks together with the proposed framework in the zero-shot evaluation of five different morphing generation techniques that are captured in three different mediums.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。