arXiv:2505.03610cs.CV2025-05TPAMI被引 7

用知识图谱生成精准提示,提升3D面具攻击检测的泛化能力。

Learning Knowledge-based Prompts for Robust 3D Mask Presentation Attack Detection

  • 基于知识图谱构建视觉语言提示,实现细粒度特征提取。
  • 在多个数据集上达到当前最优的跨场景检测性能。
  • 适合关注安全防御与多模态学习的研究者使用。

3D面具演示攻击检测对于防范面部识别系统面临日益增长的3D面具攻击至关重要。现有方法多依赖多模态特征或远程光电容积脉搏波(rPPG)信号区分真实人脸与3D面具,但面临传感器成本高和泛化能力差的问题。检测相关文本描述具有信息简洁、通用性强且获取成本低的优势。然而,视觉-语言多模态特征在3D面具攻击检测中的潜力尚未被探索。本文提出一种新型知识驱动提示学习框架,挖掘预训练视觉-语言模型中蕴含的知识,以增强3D面具攻击检测的泛化能力。具体地,将知识图谱中的实体与三元组融入提示学习过程,生成细粒度、任务特定的显式提示。针对不同输入图像可能强调不同知识图谱元素的情况,引入基于注意力机制的视觉特定知识过滤器,根据视觉上下文精炼相关元素。此外,结合因果图理论优化提示学习过程:训练时采用虚假关联消除范式,利用知识文本特征引导剔除类别无关的局部图像块,促进学习与类别相关局部区域对齐的因果提示。实验结果表明,该方法在基准数据集上实现了领先的内部与跨场景检测性能。

原文摘要 · Abstract (English)

3D mask presentation attack detection is crucial for protecting face recognition systems against the rising threat of 3D mask attacks. While most existing methods utilize multimodal features or remote photoplethysmography (rPPG) signals to distinguish between real faces and 3D masks, they face significant challenges, such as the high costs associated with multimodal sensors and limited generalization ability. Detection-related text descriptions offer concise, universal information and are cost-effective to obtain. However, the potential of vision-language multimodal features for 3D mask presentation attack detection remains unexplored. In this paper, we propose a novel knowledge-based prompt learning framework to explore the strong generalization capability of vision-language models for 3D mask presentation attack detection. Specifically, our approach incorporates entities and triples from knowledge graphs into the prompt learning process, generating fine-grained, task-specific explicit prompts that effectively harness the knowledge embedded in pre-trained vision-language models. Furthermore, considering different input images may emphasize distinct knowledge graph elements, we introduce a visual-specific knowledge filter based on an attention mechanism to refine relevant elements according to the visual context. Additionally, we leverage causal graph theory insights into the prompt learning process to further enhance the generalization ability of our method. During training, a spurious correlation elimination paradigm is employed, which removes category-irrelevant local image patches using guidance from knowledge-based text features, fostering the learning of generalized causal prompts that align with category-relevant local patches. Experimental results demonstrate that the proposed method achieves state-of-the-art intra- and cross-scenario detection performance on benchmark datasets.

3D面具检测提示学习知识图谱视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。