arXiv:2503.19982cs.CV2025-03AAAI被引 20

用语言模型增强活体检测,让单类方法更抗干扰。

SLIP: Spoof-Aware One-Class Face Anti-Spoofing with Language Image Pretraining

  • 用语言引导生成伪造线索图,模拟攻击遮挡
  • 通过提示词分离出与真假相关的特征
  • 融合真实人脸和伪造提示生成多样伪造特征

人脸反伪造(FAS)对保障人脸识别系统的安全可靠至关重要。随着视觉-语言预训练(VLP)模型的发展,近期的双类FAS方法已利用VLP优势,但单类FAS方法尚未探索此潜力。单类FAS仅从真实人脸图像中学习活体特征,以区分真实与伪造人脸,但缺乏伪造数据可能导致模型误学无关域信息(如面部内容),在新场景下性能下降。为此,本文提出一种新框架SLIP:基于真实人脸不应被攻击物体(如纸张、面具)遮挡的假设,首次引入语言引导的伪造线索图估计,通过判断是否被攻击物覆盖并生成非零线索图;进一步设计提示驱动的活体特征解耦机制,分离出与真假相关和与域相关的特征;最后,通过融合真实图像与伪造提示的潜在特征生成伪造类特征,丰富伪造特征分布。大量实验与消融研究证明,SLIP持续优于已有单类FAS方法。

原文摘要 · Abstract (English)

Face anti-spoofing (FAS) plays a pivotal role in ensuring the security and reliability of face recognition systems. With advancements in vision-language pretrained (VLP) models, recent two-class FAS techniques have leveraged the advantages of using VLP guidance, while this potential remains unexplored in one-class FAS methods. The one-class FAS focuses on learning intrinsic liveness features solely from live training images to differentiate between live and spoof faces. However, the lack of spoof training data can lead one-class FAS models to inadvertently incorporate domain information irrelevant to the live/spoof distinction (e.g., facial content), causing performance degradation when tested with a new application domain. To address this issue, we propose a novel framework called Spoof-aware one-class face anti-spoofing with Language Image Pretraining (SLIP). Given that live faces should ideally not be obscured by any spoof-attack-related objects (e.g., paper, or masks) and are assumed to yield zero spoof cue maps, we first propose an effective language-guided spoof cue map estimation to enhance one-class FAS models by simulating whether the underlying faces are covered by attack-related objects and generating corresponding nonzero spoof cue maps. Next, we introduce a novel prompt-driven liveness feature disentanglement to alleviate live/spoof-irrelative domain variations by disentangling live/spoof-relevant and domain-dependent information. Finally, we design an effective augmentation strategy by fusing latent features from live images and spoof prompts to generate spoof-like image features and thus diversify latent spoof features to facilitate the learning of one-class FAS. Our extensive experiments and ablation studies support that SLIP consistently outperforms previous one-class FAS methods.

活体检测语言模型单类学习反伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。