arXiv:2505.24214cs.CVcs.AI2025-05被引 15

零样本评估大模型在六类生物特征任务中的表现,无需微调即达高精度。

Benchmarking Foundation Models for Zero-Shot Biometric Tasks

  • 用41个预训练多模态模型零样本测试人脸与虹膜任务
  • 无微调下人脸识别准确率达96.77%,虹膜识别达97.55%
  • 简单分类头可有效实现换脸检测与攻击识别,适合安全应用

基础模型(尤其是视觉-语言模型和多模态大语言模型)的兴起重新定义了人工智能边界,实现了跨任务的出色泛化能力。然而其在生物特征识别与分析中的潜力仍待挖掘。本文构建了一个综合性基准,评估41个公开可用的VLM与MLLM在六类生物特征任务上的零样本与少样本性能,涵盖人脸与虹膜模态:人脸验证、软生物特征属性预测(性别与种族)、虹膜识别、活体攻击检测(PAD)及人脸篡改检测(变形与深度伪造)。实验表明,这些模型的嵌入表示可在无需微调的情况下完成多种生物特征任务。例如,在无微调条件下,于LFW数据集上人脸验证的真匹配率(TMR)达96.77%(假匹配率FMR为1%);在IITD-R-Full数据集上,虹膜识别的TMR达97.55%(FMR为1%)。此外,通过添加简单分类头,可实现对人脸深度伪造的检测、虹膜活体攻击识别,以及从人脸中提取性别、种族等软生物特征属性,并取得合理高精度。本工作再次印证了预训练模型在实现通用人工智能愿景中的巨大潜力。

原文摘要 · Abstract (English)

The advent of foundation models, particularly Vision-Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), has redefined the frontiers of artificial intelligence, enabling remarkable generalization across diverse tasks with minimal or no supervision. Yet, their potential in biometric recognition and analysis remains relatively underexplored. In this work, we introduce a comprehensive benchmark that evaluates the zero-shot and few-shot performance of state-of-the-art publicly available VLMs and MLLMs across six biometric tasks spanning the face and iris modalities: face verification, soft biometric attribute prediction (gender and race), iris recognition, presentation attack detection (PAD), and face manipulation detection (morphs and deepfakes). A total of 41 VLMs were used in this evaluation. Experiments show that embeddings from these foundation models can be used for diverse biometric tasks with varying degrees of success. For example, in the case of face verification, a True Match Rate (TMR) of 96.77 percent was obtained at a False Match Rate (FMR) of 1 percent on the Labeled Face in the Wild (LFW) dataset, without any fine-tuning. In the case of iris recognition, the TMR at 1 percent FMR on the IITD-R-Full dataset was 97.55 percent without any fine-tuning. Further, we show that applying a simple classifier head to these embeddings can help perform DeepFake detection for faces, Presentation Attack Detection (PAD) for irides, and extract soft biometric attributes like gender and ethnicity from faces with reasonably high accuracy. This work reiterates the potential of pretrained models in achieving the long-term vision of Artificial General Intelligence.

生物特征识别零样本学习大模型应用多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。