用大模型提升虹膜活体检测,小样本下效果超越现有方法
Towards Iris Presentation Attack Detection with Foundation Models
- 基于DinoV2和VisualOpenClip大模型,微调轻量分类头
- 小样本下性能超过当前最先进深度学习方法
- 适合数据稀缺的虹膜活体检测场景
由于在大规模数据集上训练,基础模型具备强大的泛化能力,这在近红外虹膜活体攻击检测(NIR Iris PAD)中尤为有用。该领域数据库受限于受试者数量和攻击手段多样性,且真品与攻击图像通常来自不同个体,缺乏对应关系。本文探索了基于DinoV2和VisualOpenClip两个基础模型的虹膜PAD方法。结果表明,使用小型神经网络作为头部对模型进行微调,性能超越当前基于深度学习的最先进方法。然而,若同时拥有真品和攻击图像,从零开始训练的系统仍能取得更优结果。
原文摘要 · Abstract (English)
Foundation models are becoming increasingly popular due to their strong generalization capabilities resulting from being trained on huge datasets. These generalization capabilities are attractive in areas such as NIR Iris Presentation Attack Detection (PAD), in which databases are limited in the number of subjects and diversity of attack instruments, and there is no correspondence between the bona fide and attack images because, most of the time, they do not belong to the same subjects. This work explores an iris PAD approach based on two foundation models, DinoV2 and VisualOpenClip. The results show that fine-tuning prediction with a small neural network as head overpasses the state-of-the-art performance based on deep learning approaches. However, systems trained from scratch have still reached better results if bona fide and attack images are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。