用视觉特征组合方式,让模型自动识别未知伪造人脸
Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

- 将伪造证据拆解为可复用的微小视觉单元,动态组合成新攻击识别提示
- 在9个开放域测试中达到当前最优性能,对未见攻击仍保持高准确率
- 无需预设语义标签,通过共享参数自适应学习视觉特征的组合规律
开放世界人脸识别反伪造需应对分布与语义双重变化:源域与目标域成像条件不同,且目标域包含训练时未见的攻击类型。现有基于提示的方法多依赖类别语义或语言引导,虽能建模高层概念,但难以捕捉未见攻击中不断演化的细粒度、空间异质的取证证据。我们假设:许多未见攻击可由重复出现的视觉线索新组合构成。为此提出一种完全在视觉特征空间运行的组合式取证视觉提示学习框架。基于冻结的ViT视觉基础模型,利用块感知注意力将共享的可学习微取证原型细化为图像块生成的局部取证证据单元。类别特定的全局上下文提示提供输入相关路由权重,自适应选择并组合这些原型,形成用于真实/伪造判别的组合式视觉提示。原型不预设语义含义,其专属性与复用性由跨类别的共享参数化与联合优化自然产生。在九个开放世界协议上的大量实验表明,该方法达到最先进性能,具备强跨域泛化能力,并对未见攻击有良好适应性。
原文摘要 · Abstract (English)
Open-world face anti-spoofing must address both covariate and semantic shifts: source and target domains differ in imaging conditions, while target domains contain diverse attack types absent from training. Existing prompt-based approaches often express spoofing through category semantics or language guidance, which is effective for modeling high-level concepts but is less suited to explicitly capturing the evolving fine-grained and spatially heterogeneous forensic evidence of unseen attacks. Motivated by the hypothesis that many unseen attacks can be characterized by new combinations of recurring visual cues, we propose a compositional forensic visual prompt learning framework that operates entirely in the visual feature space. Built on a frozen ViT-based vision foundation model, the framework employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches. Class-specific global contextual prompts then provide input-dependent routing weights that adaptively select and compose these primitives into compositional forensic visual prompts for real/spoof discrimination. The primitives are not assigned predefined semantic meanings; instead, their specialization and reuse emerge from shared parameterization and joint optimization across categories. Extensive experiments on nine open-world protocols demonstrate state-of-the-art performance, strong cross-domain generalization, and robust adaptation to unseen attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。