评测大模型在跨模态人脸识别中的表现,发现其仍远不如传统系统。
Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition
- 用多模态大模型对比不同传感器下的面部识别效果
- 在可见光到红外等跨谱条件下,识别准确率显著低于传统方法
- 适合关注大模型在生物识别中应用潜力的研究者
多模态大语言模型(MLLMs)在众多视觉-语言任务中表现出色,引发了对其在生物识别中应用的兴趣。本文系统评估了当前先进的MLLMs在异构人脸识别(HFR)中的性能,其中注册图像与测试图像来自不同传感模态,包括可见光(VIS)、近红外(NIR)、短波红外(SWIR)和热成像相机。我们在多个跨模态场景(如VIS-NIR、VIS-SWIR、VIS-THERMAL)中对多个开源MLLMs进行了基准测试。通过生物识别协议和多种指标(包括获取率、等错误率EER、真实接受率TAR)评估识别性能。结果表明,尽管MLLMs取得进展,但在挑战性跨光谱条件下,其性能与经典人脸识别系统存在显著差距。研究揭示了当前MLLMs在异构人脸识别中的局限性,也强调了在部署前进行严格生物识别评估的重要性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance on a wide range of vision-language tasks, raising interest in their potential use for biometric applications. In this paper, we conduct a systematic evaluation of state-of-the-art MLLMs for heterogeneous face recognition (HFR), where enrollment and probe images are from different sensing modalities, including visual (VIS), near infrared (NIR), short-wave infrared (SWIR), and thermal camera. We benchmark multiple open-source MLLMs across several cross-modality scenarios, including VIS-NIR, VIS-SWIR, and VIS-THERMAL face recognition. The recognition performance of MLLMs is evaluated using biometric protocols and based on different metrics, including Acquire Rate, Equal Error Rate (EER), and True Accept Rate (TAR). Our results reveal substantial performance gaps between MLLMs and classical face recognition systems, particularly under challenging cross-spectral conditions, in spite of recent advances in MLLMs. Our findings highlight the limitations of current MLLMs for HFR and also the importance of rigorous biometric evaluation when considering their deployment in face recognition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。