arXiv:2607.01303cs.CVcs.AI2026-07中稿 · IEEE Transactions …被引 1

用视觉概念引导提示,提升人脸识别抗欺骗能力

CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection

论文配图:CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection
图 1 · 摘自论文原文
  • 通过XAI自动发现攻击相关视觉概念,生成局部指导热图
  • 提示词注入机制使模型捕捉跨域通用攻击特征,抑制数据偏差
  • 在9个数据集上实现领先跨域检测效果,适合安全敏感场景

活体攻击检测(PAD)是防范打印照片、重放视频和3D面具等欺骗攻击的关键防线。尽管进展显著,现有模型在未见领域间泛化能力仍不足,受限于传感器、光照及攻击材料差异。近期视觉-语言模型虽具强泛化性,但其提示词通常依赖类别标签优化,难以显式对齐细粒度攻击相关语义,导致表征过拟合特定域特征而非捕获可迁移的攻击线索。为此,本文提出概念引导提示的活体攻击检测框架CPG-PAD,将模型级概念引导融入提示学习过程。具体设计了视觉概念增强(VCE)模块,利用可解释AI技术自动发现与PAD相关的视觉概念,并生成关联热图以提供细粒度局部指导。基于热图,提示词概念注入(PCI)机制通过视觉提示解码器(VPD)和概念映射损失,将概念融入提示空间,使提示与模型内部概念空间对齐。该设计有效捕捉可迁移的、域不变的攻击线索,同时抑制数据集特异性偏差。在九个基准数据集上的大量实验表明,无论多源、少源还是单源设置下,CPG-PAD均持续达到最优跨域性能。

原文摘要 · Abstract (English)

Presentation Attack Detection (PAD) serves as a crucial safeguard for face recognition systems against presentation attacks such as printed photos, replayed videos, and 3D masks. Despite significant progress, existing PAD models still struggle to generalize across unseen domains due to variations in sensors, lighting, and attack materials. Recent Vision-Language Models (VLMs) have shown strong generalization ability, yet their applications in PAD remain limited because learned prompts, typically optimized under class-label supervision, fail to explicitly align with fine-grained attack-relevant visual semantics. As a result, the learned representations often overfit domain-specific artifacts instead of capturing transferable attack cues. To address this, we propose Concept-Informed Prompts Guided Presentation Attack Detection (CPG-PAD), a framework that introduces model-level concept guidance into the prompt learning process. Specifically, we design a Visual Concept-driven Enhancement (VCE) module that employs eXplainable AI (XAI) techniques to automatically discover PAD-relevant visual concepts and generate concept-associated heatmaps providing localized fine-grained guidance. Guided by these heatmaps, a Prompt-based Concept Injection (PCI) mechanism integrates these concepts into the prompt space through a Visual-Prompt Decoder (VPD) and a concept-mapping loss, enabling prompts to align with the model's internal concept space. This design enables CPG-PAD to capture generalizable and domain-invariant attack cues while effectively suppressing dataset-specific biases. Extensive experiments across nine benchmark datasets demonstrate that CPG-PAD consistently achieves state-of-the-art cross-domain performance under multi-source, limited-source, and single-source settings.

活体检测视觉概念提示学习跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。