arXiv:2501.01243cs.CVcs.AI2025-01NeurIPS被引 13

构建首个面向多模态助手的人脸与人体理解评测基准。

Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants

  • 提出三级能力分类体系,系统梳理人脸人体理解维度。
  • 构建含3600个问题的双语评测集,支持中英文测试。
  • 揭示主流模型在位置敏感性与思维链提示下的表现差异。

人脸与人体是社交互动的关键元素,广泛存在于日常图像与视频中。深入理解人脸与人体,有助于提升多模态助手的响应质量与应用范围。然而,当前多模态助手领域缺乏对人脸与人体理解能力的全面、科学评估。本文首先提出包含三个层级能力的层次化能力分类体系;基于该体系,从公开数据集中收集图像与标注,并构建半自动数据流水线生成评测问题。最终形成的Face-Human-Bench包含开发集与测试集各1800个问题,支持中英文双语。我们在25个主流多模态大模型(MLLMs)上进行评测,重点分析能力间的相关性、目标相对位置对性能的影响,以及思维链(CoT)提示的作用。同时探讨了哪些能力需由专用模型补充。数据集与评估代码已公开于https://face-human-bench.github.io。

原文摘要 · Abstract (English)

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality and broadened application scope. Currently, the multi-modal assistant community lacks a comprehensive and scientific evaluation of face and human understanding abilities. In this paper, we first propose a hierarchical ability taxonomy that includes three levels of abilities. Then, based on this taxonomy, we collect images and annotations from publicly available datasets in the face and human community and build a semi-automatic data pipeline to produce problems for the new benchmark. Finally, the obtained Face-Human-Bench includes a development set and a test set, each with 1800 problems, supporting both English and Chinese. We conduct evaluations over 25 mainstream multi-modal large language models (MLLMs) with our Face-Human-Bench, focusing on the correlation between abilities, the impact of the relative position of targets on performance, and the impact of Chain of Thought (CoT) prompting on performance. We also explore which abilities of MLLMs need to be supplemented by specialist models. The dataset and evaluation code have been made publicly available at https://face-human-bench.github.io.

多模态评测基准人脸理解人体理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。