arXiv:2606.18510cs.CVcs.CR2026-06

Vision Transformer比传统CNN在人脸伪造检测中更公平,尤其对未见过的族裔群体表现更好。

Architectural Bias in Face Presentation Attack Detection: A Comparative Study of Vision Transformers and Convolutional Neural Networks

论文配图:Architectural Bias in Face Presentation Attack Detection: A Comparative Study of Vision Transformers and Convolutional Neural Networks
图 1 · 摘自论文原文
  • 用ViT和CNN对比测试,发现预训练ViT泛化能力更强
  • 在非洲与东亚人群间误检率差距缩小至0.13%,降低83%
  • 对未知族裔群体的误报率仅为CNN的1/3.6,适合跨种族应用

人脸伪造攻击检测(PAD)是生物识别认证的关键安全层,但现有方法在不同族裔间存在系统性性能差异,尤其对深肤色个体不利。本文通过实证比较了视觉变压器(ViT)与卷积神经网络(CNN)在减少此类偏差方面的潜力。实验基于跨族裔人脸防伪数据集CeFA,评估了三种模型:从头训练的Multimodal ViT-Tiny、ResNet18基准模型,以及在CeFA上微调的预训练DeiT-S,覆盖非洲、东亚及零样本中央亚族裔群体。DeiT-S取得最高总体准确率97.27%和最低等错误率(EER)0.86%,优于ResNet18的90.15%。在公平性方面,德伊特-S将非洲与东亚受试者间的跨族裔ACER差距降至0.13%,较此前基于LBP的工作(0.75%)降低83%。最显著的是,当面对零样本中央亚族裔时,ResNet18的BPCER高达10.44%,而DeiT-S仅2.89%,展现出3.6倍的泛化优势。结果表明,预训练的视觉变压器在提高精度、缩小族裔性能差距和增强跨族裔泛化方面具有显著优势,提示架构设计可能影响跨族裔公平性。

原文摘要 · Abstract (English)

Face Presentation Attack Detection (PAD) systems constitute a critical security layer in biometric authentication; however, existing approaches exhibit systematic performance disparities across demographic groups, disproportionately affecting individuals with darker skin tones. This paper presents a comparative empirical investigation of whether Vision Transformer architectures reduce demographic bias in face PAD systems relative to convolutional baselines. Experiments are conducted on the CASIA-SURF Cross-Ethnicity Face Anti-Spoofing (CeFA) dataset. Three architectures are evaluated: a Multimodal ViT-Tiny trained from scratch, a ResNet18 CNN baseline, and a pretrained DeiT-S fine-tuned on CeFA across African, East Asian, and zero-shot Central Asian demographic groups. DeiT-S achieves the highest overall accuracy of 97.27% and the lowest EER of 0.86%, outperforming ResNet18 at 90.15% accuracy. In terms of fairness, DeiT-S reduces the inter-ethnic ACER gap between African and East Asian subjects to 0.13%, compared to 0.75% reported in an LBP-based work [6], representing an 83% reduction. Most notably, while ResNet18 records a BPCER of 10.44% on zero-shot Central Asian subjects, DeiT-S maintains 2.89% on the same unseen group, demonstrating a 3.6x generalization advantage. These results suggest that pretrained Vision Transformers achieve superior PAD accuracy, produce smaller demographic performance gaps, and generalize more equitably across unseen demographic groups, indicating that cross-demographic fairness in PAD may partly be influenced by architectural design.

人脸反欺诈视觉变压器公平性跨族裔泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。