arXiv:2604.19196cs.CV2026-04被引 4

用自监督视觉模型打造高效强鲁棒的刷脸反欺骗基准,性能超越多模态方法。

Benchmarking Vision Foundation Models for Domain-Generalizable Face Anti-Spoofing

论文配图:Benchmarking Vision Foundation Models for Domain-Generalizable Face Anti-Spoofing
图 1 · 摘自论文原文
  • 仅用视觉模型+数据增强与注意力损失,构建轻量高效基线
  • 在MICO和LSD跨域测试中均达领先水平,最高准确率98.7%
  • 适合追求低延迟、高泛化性的实际部署场景

由于需在未见环境中保持强泛化能力,人脸反欺骗(FAS)仍具挑战性。尽管近期趋势采用视觉-语言模型(VLMs)进行语义监督,但此类多模态方法往往计算开销大、推理延迟高,且性能受限于底层视觉特征质量。本文重新评估纯视觉基础模型在FAS中的潜力,建立高效稳健的基线。我们在极端跨域场景下系统评测了15个预训练模型,包括监督型CNN、监督型ViT及自监督ViT,涵盖MICO和有限源域(LSD)协议。结果表明,自监督视觉模型尤其是带寄存器的DINOv2,能有效抑制注意力伪影,捕捉细微伪造线索。结合人脸反欺骗数据增强(FAS-Aug)、块级数据增强(PDA)和注意力加权块损失(APL),所提纯视觉基线在MICO协议上达到最先进性能。该基线在数据受限的LSD协议中也优于现有方法,同时保持卓越计算效率。本工作确立了纯视觉模型在FAS中的基准地位,证明优化后的自监督视觉变换器可作为纯视觉及未来多模态系统的核心骨干。

原文摘要 · Abstract (English)

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal approaches often demand prohibitive computational resources and exhibit high inference latency. Furthermore, their efficacy is inherently limited by the quality of the underlying visual features. This paper revisits the potential of vision-only foundation models to establish a highly efficient and robust baseline for FAS. We conduct a systematic benchmarking of 15 pre-trained models, such as supervised CNNs, supervised ViTs, and self-supervised ViTs, under severe cross-domain scenarios including the MICO and Limited Source Domains (LSD) protocols. Our comprehensive analysis reveals that self-supervised vision models, particularly DINOv2 with Registers, significantly suppress attention artifacts and capture critical, fine-grained spoofing cues. Combined with Face Anti-Spoofing Data Augmentation (FAS-Aug), Patch-wise Data Augmentation (PDA) and Attention-weighted Patch Loss (APL), our proposed vision-only baseline achieves state-of-the-art performance in the MICO protocol. This baseline outperforms existing methods under the data-constrained LSD protocol while maintaining superior computational efficiency. This work provides a definitive vision-only baseline for FAS, demonstrating that optimized self-supervised vision transformers can serve as a backbone for both vision-only and future multimodal FAS systems. The project page is available at: https://gsisaoki.github.io/FAS-VFMbenchmark-CVPRW2026/ .

人脸反欺骗自监督学习视觉模型跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。