构建多视角细粒度人脸属性数据集,评测大模型识脸能力
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
- 设计五视角三层次210+属性结构,生成近7.4万组视觉问答对
- 自研模型Face-LLaVA仅用少量数据即超越多数开源模型
- 适合作为评估人脸识别大模型性能的基准数据集
多模态大语言模型在多种任务中展现出强大能力,但其在人脸感知方面的有效评估仍不充分。为此,我们提出FaceBench,一个包含分层多视角、多层级属性的数据集,旨在全面评估多模态大语言模型的人脸感知能力。首先构建包含五个视角、最多三层属性的层级化人脸属性体系,共涵盖超过210个属性和700个属性值。基于该结构,构建了49,919个用于评估的视觉问答对和23,841个用于微调的问答对。此外,我们通过使用该数据集训练出一个鲁棒的人脸感知多模态大模型基线——Face-LLaVA。在主流多模态大模型及Face-LLaVA上开展大量实验,测试其人脸感知能力,并与人类表现进行对比。结果表明,现有模型在细粒度人脸属性理解方面仍不理想;而我们的Face-LLaVA在少量训练数据下显著优于多数开源模型,且接近GPT-4o和Gemini等商用模型水平。数据集将发布于https://github.com/CVI-SZU/FaceBench。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level attributes specifically designed to assess the comprehensive face perception abilities of MLLMs. Initially, we construct a hierarchical facial attribute structure, which encompasses five views with up to three levels of attributes, totaling over 210 attributes and 700 attribute values. Based on the structure, the proposed FaceBench consists of 49,919 visual question-answering (VQA) pairs for evaluation and 23,841 pairs for fine-tuning. Moreover, we further develop a robust face perception MLLM baseline, Face-LLaVA, by training with our proposed face VQA data. Extensive experiments on various mainstream MLLMs and Face-LLaVA are conducted to test their face perception ability, with results also compared against human performance. The results reveal that, the existing MLLMs are far from satisfactory in understanding the fine-grained facial attributes, while our Face-LLaVA significantly outperforms existing open-source models with a small amount of training data and is comparable to commercial ones like GPT-4o and Gemini. The dataset will be released at https://github.com/CVI-SZU/FaceBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。