arXiv:2504.04470cs.CV2025-04被引 29

通过内容感知提示生成,提升人脸识别反欺骗的跨域泛化能力。

Domain Generalization for Face Anti-spoofing via Content-aware Composite Prompt Engineering

  • 用实例级复合提示替代类别级提示,融合固定模板与可学习部分。
  • 在多个跨域测试中达到当前最佳性能,显著提升对细微伪造痕迹的识别。
  • 适合需要高泛化能力的安防、金融场景下的活体检测应用。

人脸反欺骗中的领域泛化(DG)面临域特定信号干扰细微伪造线索的问题。现有基于CLIP的算法通过调整视觉分类器权重缓解此问题,但存在两大缺陷:(1) 真/伪人脸等类别对CLIP模型无语义,难以学习准确描述;(2) 单一提示无法覆盖多种伪造类型。为此,本文提出内容感知复合提示工程(CCPE),生成实例级复合提示,包含固定模板与可学习提示。具体地,从两个分支构建内容感知提示:(1) 固有内容提示利用基于指令的大语言模型(LLM)迁移丰富知识;(2) 可学习内容提示通过Q-Former隐式提取最有效视觉信息。此外,设计跨模态引导模块(CGM),动态调节单模态特征以实现更优融合。在多个跨域实验中验证了有效性,取得当前最优(SOTA)结果。

原文摘要 · Abstract (English)

The challenge of Domain Generalization (DG) in Face Anti-Spoofing (FAS) is the significant interference of domain-specific signals on subtle spoofing clues. Recently, some CLIP-based algorithms have been developed to alleviate this interference by adjusting the weights of visual classifiers. However, our analysis of this class-wise prompt engineering suffers from two shortcomings for DG FAS: (1) The categories of facial categories, such as real or spoof, have no semantics for the CLIP model, making it difficult to learn accurate category descriptions. (2) A single form of prompt cannot portray the various types of spoofing. In this work, instead of class-wise prompts, we propose a novel Content-aware Composite Prompt Engineering (CCPE) that generates instance-wise composite prompts, including both fixed template and learnable prompts. Specifically, our CCPE constructs content-aware prompts from two branches: (1) Inherent content prompt explicitly benefits from abundant transferred knowledge from the instruction-based Large Language Model (LLM). (2) Learnable content prompts implicitly extract the most informative visual content via Q-Former. Moreover, we design a Cross-Modal Guidance Module (CGM) that dynamically adjusts unimodal features for fusion to achieve better generalized FAS. Finally, our CCPE has been validated for its effectiveness in multiple cross-domain experiments and achieves state-of-the-art (SOTA) results.

人脸反欺骗领域泛化提示工程多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。