arXiv:2607.20900cs.CV2026-07中稿 · IJCB2026

用视觉提示调优实现物理与数字人脸欺骗的统一检测

DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing

论文配图:DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing
图 1 · 摘自论文原文
  • 通过分层视觉提示注入,动态分离不同伪造特征
  • 在UniAttackData上超越现有视觉语言模型方法
  • 仅用视觉模态即达顶尖性能,无需文本监督

随着伪造攻击形式日益多样,亟需能同时识别物理和数字人脸欺骗的统一反伪造模型。现有视觉-语言模型虽具强泛化能力,但依赖复杂的多模态融合与外部文本编码器。本文提出DINO-VPT,一种轻量级纯视觉框架,采用分层视觉提示调优。通过提示路由网络(PRN)根据输入特征动态注入提示,有效解耦多种伪造痕迹,无需多模态融合。在UniAttackData基准上的评估表明,DINO-VPT性能优于当前最优的VLM方法。结果表明,结构合理的纯视觉架构可在无需多模态监督的情况下实现统一人脸反伪造的顶尖表现。

原文摘要 · Abstract (English)

With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physical and digital threats. While existing Vision-Language Models (VLMs) demonstrate high generalization in this context, they heavily rely on complex multimodal fusion and external text encoders. In this paper, we propose DINO-VPT, a lightweight, vision-only framework leveraging hierarchical visual prompt tuning. By dynamically injecting prompts conditioned on input features via a Prompt Routing Network (PRN), our method effectively disentangles diverse spoofing artifacts without requiring multimodal fusion. Evaluations on the UniAttackData benchmark demonstrate that DINO-VPT achieves higher accuracy than state-of-the-art VLM-based methods. Our results indicate that a properly structured vision-only architecture can achieve state-of-the-art performance in unified FAS without the need for multimodal supervision.

人脸识别反欺骗视觉提示多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。