arXiv:2507.05575cs.CV2025-07被引 4

通过跨模态特征过渡提升多模态人脸反欺骗的鲁棒性

Multi-Modal Face Anti-Spoofing via Cross-Modal Feature Transitions

  • 利用活体样本间跨模态特征过渡的一致性构建通用特征空间
  • 通过区分活体与伪造间特征过渡不一致性,有效检测分布外攻击
  • 从RGB模态学习红外与深度辅助特征,应对模态缺失问题

多模态人脸反欺骗(FAS)通过融合RGB、红外(IR)和深度图像等多源信息提取活体线索,提升生物识别系统的鲁棒性。然而,由于不同模态由异构传感器采集且受环境差异影响,多模态FAS在训练与测试域间存在显著分布差异。此外,推理阶段常面临部分模态缺失的问题。本文提出交叉模态过渡引导网络(CTNet),核心思想是:同一模态内活体样本间视觉差异较小,而跨模态特征过渡在活体样本中更一致,且远高于活体与伪造间的过渡一致性。为此,我们首先学习活体样本间的跨模态特征过渡以构建泛化特征空间;其次,学习活体与伪造间过渡不一致性以检测分布外(OOD)攻击。为应对模态缺失,还提出从RGB模态中学习互补的红外(IR)与深度特征作为辅助模态。大量实验表明,所提方法在多数协议下优于现有两分类多模态FAS方法。

原文摘要 · Abstract (English)

Multi-modal face anti-spoofing (FAS) aims to detect genuine human presence by extracting discriminative liveness cues from multiple modalities, such as RGB, infrared (IR), and depth images, to enhance the robustness of biometric authentication systems. However, because data from different modalities are typically captured by various camera sensors and under diverse environmental conditions, multi-modal FAS often exhibits significantly greater distribution discrepancies across training and testing domains compared to single-modal FAS. Furthermore, during the inference stage, multi-modal FAS confronts even greater challenges when one or more modalities are unavailable or inaccessible. In this paper, we propose a novel Cross-modal Transition-guided Network (CTNet) to tackle the challenges in the multi-modal FAS task. Our motivation stems from that, within a single modality, the visual differences between live faces are typically much smaller than those of spoof faces. Additionally, feature transitions across modalities are more consistent for the live class compared to those between live and spoof classes. Upon this insight, we first propose learning consistent cross-modal feature transitions among live samples to construct a generalized feature space. Next, we introduce learning the inconsistent cross-modal feature transitions between live and spoof samples to effectively detect out-of-distribution (OOD) attacks during inference. To further address the issue of missing modalities, we propose learning complementary infrared (IR) and depth features from the RGB modality as auxiliary modalities. Extensive experiments demonstrate that the proposed CTNet outperforms previous two-class multi-modal FAS methods across most protocols.

人脸反欺骗多模态特征过渡鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。