通过跨模态语义对齐提升人脸伪造检测泛化能力
XSA-MAD: Cross-modal Semantic Alignment for Morphing Attack Detection

- 用四个可解释属性构建文本表征,对齐图像语义空间
- 在GAN生成伪造上实现2.92%等错误率,优于现有方法
- 适合需要跨生成方式泛化的伪造检测场景
人脸伪装攻击严重威胁人脸识别系统。现有基于图像的伪装攻击检测(MAD)方法因仅依赖视觉线索,难以泛化到未见生成技术。本文提出XSA-MAD,一种基于CLIP的多模态框架,显式建模真实与伪造人脸间的语义不一致。将伪装概念分解为身份、面部几何、纹理和一致性四类可解释属性,并生成结构化、属性感知的文本表征。图像编码器逐步对齐该判别性文本空间,形成统一的语义表示,捕捉生成无关且概念级的真实与伪造差异。在SMDD训练后,于MAD22和MorDIFF数据集上的实验表明,该方法对多种伪装原理具有强泛化能力。尤其在基于GAN的伪装上,达到2.92%等错误率,且在高保真生成攻击下持续优于现有方法。
原文摘要 · Abstract (English)
Morphing attacks pose a serious threat to face recognition systems. However, existing image-based morphing attack detection (MAD) methods often generalize poorly to unseen generation techniques because they rely solely on visual cues. We propose XSA-MAD, a CLIP-based multimodal framework that explicitly models semantic inconsistencies between bona-fide and morphed faces. Morphing concepts are decomposed into four interpretable attributes, including identity, facial geometry, texture, and consistency, and are encoded as structured and attribute-aware textual representations. The image encoder is progressively aligned with this discriminative textual space, resulting in a unified semantic representation that captures generation-invariant and concept-level discrepancies between bona-fide and morph images. Experiments on MAD22 and MorDIFF, following training on SMDD, demonstrate strong generalization across diverse morphing principles. In particular, XSA-MAD achieves an equal error rate of 2.92% on GAN-based morphs and consistently outperforms existing methods under high-fidelity generative attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。