评测开源多模态大模型在人脸识别上的表现,发现其通用性好但精度不如专用模型。
Benchmarking Multimodal Large Language Models for Face Recognition
- 在多个标准数据集上系统评估主流多模态大模型的人脸识别能力。
- 零样本下模型在高精度场景中性能落后于专用人脸识别模型。
- 为后续模型设计提供基准参考,适合关注跨模态识别的研究者。
多模态大语言模型(MLLM)在多种视觉-语言任务中表现出色,但在人脸识别领域的潜力尚未充分探索。特别是开源的MLLM在与现有专用人脸识别模型对比时,缺乏在标准基准上使用相似协议的系统评估。本文针对LFW、CALFW、CPLFW、CFP、AgeDB和RFW等多个数据集,对当前最先进的多模态大模型进行了系统性人脸识别基准测试。实验结果表明,尽管MLLM能捕捉丰富语义信息,适用于相关任务,但在零样本应用的高精度识别场景中仍显著落后于专用模型。本研究为推进基于MLLM的人脸识别提供了基准支持,并为下一代更精准、泛化能力更强的模型设计提供了洞见。项目代码已公开。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved remarkable performance across diverse vision-and-language tasks. However, their potential in face recognition remains underexplored. In particular, the performance of open-source MLLMs needs to be evaluated and compared with existing face recognition models on standard benchmarks with similar protocol. In this work, we present a systematic benchmark of state-of-the-art MLLMs for face recognition on several face recognition datasets, including LFW, CALFW, CPLFW, CFP, AgeDB and RFW. Experimental results reveal that while MLLMs capture rich semantic cues useful for face-related tasks, they lag behind specialized models in high-precision recognition scenarios in zero-shot applications. This benchmark provides a foundation for advancing MLLM-based face recognition, offering insights for the design of next-generation models with higher accuracy and generalization. The source code of our benchmark is publicly available in the project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。