arXiv:2503.22936cs.CV2025-03被引 3

通过三类新训练策略提升面部活体检测模型性能

Enhancing Learnable Descriptive Convolutional Vision Transformer for Face Anti-Spoofing

  • 引入区域注意力监督,精准捕捉活体细微特征
  • 生成挑战性数据增强特征判别力,准确率显著提升
  • 跨域特征过渡挖掘策略,提升模型泛化能力

面部活体检测(FAS)依赖于识别真实与伪造人脸的判别性特征以应对攻击。我们此前提出的LDCformer将可学习描述性卷积(LDC)融入视觉变换器(ViT),有效建模局部描述性特征的长程依赖关系。本文提出三种新型训练策略,显著增强LDCformer的特征表征能力:第一,双注意力监督,利用区域活体/伪造注意力引导细粒度特征学习;第二,自挑战监督,通过生成高难度训练样本提升特征判别性;第三,过渡三元组挖掘策略,在缩小跨域差距的同时保持真实与伪造特征间的过渡关系,增强模型域泛化能力。大量实验表明,联合三种策略的LDCformer超越现有方法。

原文摘要 · Abstract (English)

Face anti-spoofing (FAS) heavily relies on identifying live/spoof discriminative features to counter face presentation attacks. Recently, we proposed LDCformer to successfully incorporate the Learnable Descriptive Convolution (LDC) into ViT, to model long-range dependency of locally descriptive features for FAS. In this paper, we propose three novel training strategies to effectively enhance the training of LDCformer to largely boost its feature characterization capability. The first strategy, dual-attention supervision, is developed to learn fine-grained liveness features guided by regional live/spoof attentions. The second strategy, self-challenging supervision, is designed to enhance the discriminability of the features by generating challenging training data. In addition, we propose a third training strategy, transitional triplet mining strategy, through narrowing the cross-domain gap while maintaining the transitional relationship between live and spoof features, to enlarge the domain-generalization capability of LDCformer. Extensive experiments show that LDCformer under joint supervision of the three novel training strategies outperforms previous methods.

活体检测视觉变压器特征学习深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。