首个系统评估多模态大模型人脸理解能力的基准测试
FaceXBench: Evaluating Multimodal LLMs on Face Understanding
- 构建包含5000题的跨数据集人脸理解评测集
- 26个开源及2个闭源模型表现均不理想,尤其在偏见与公平性任务上
- 适合关注人脸识别、伦理安全与多模态模型评估的研究者
多模态大语言模型(MLLMs)在多种任务中展现出强大问题解决能力,但其在人脸理解方面的能力尚未得到系统研究。为此,我们提出了FaceXBench,一个全面的基准测试,用于评估MLLMs在复杂人脸理解任务中的表现。该基准包含从25个公开数据集和新创建的FaceXAPI数据集衍生出的5000个多模态选择题,涵盖6大类共14项任务,评估内容包括偏见与公平性、人脸认证、识别、分析、定位及工具检索。我们对26个开源模型和2个专有模型进行了广泛评估,考察了零样本、上下文任务描述和思维链提示三种设置。详细分析表明,当前主流模型(如GPT-4o、GeminiPro 1.5)在复杂人脸理解任务中仍有显著提升空间。我们相信FaceXBench将成为推动具备高级人脸理解能力的MLLMs发展的关键资源。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) demonstrate impressive problem-solving abilities across a wide range of tasks and domains. However, their capacity for face understanding has not been systematically studied. To address this gap, we introduce FaceXBench, a comprehensive benchmark designed to evaluate MLLMs on complex face understanding tasks. FaceXBench includes 5,000 multimodal multiple-choice questions derived from 25 public datasets and a newly created dataset, FaceXAPI. These questions cover 14 tasks across 6 broad categories, assessing MLLMs' face understanding abilities in bias and fairness, face authentication, recognition, analysis, localization and tool retrieval. Using FaceXBench, we conduct an extensive evaluation of 26 open-source MLLMs alongside 2 proprietary models, revealing the unique challenges in complex face understanding tasks. We analyze the models across three evaluation settings: zero-shot, in-context task description, and chain-of-thought prompting. Our detailed analysis reveals that current MLLMs, including advanced models like GPT-4o, and GeminiPro 1.5, show significant room for improvement. We believe FaceXBench will be a crucial resource for developing MLLMs equipped to perform sophisticated face understanding. Code: https://github.com/Kartik-3004/facexbench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。