arXiv:2509.26165cs.CV2025-09被引 4

构建首个全面评估人本多模态大模型的基准,覆盖43个子领域。

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

  • 设计包含43个子领域的多样化人像场景,覆盖多视角、多任务
  • 提供19,945对真实图像问答,涵盖从感知到因果推理的8个维度
  • 融合自动化与人工标注,支持复杂问题如多人多图联合理解

多模态大语言模型在视觉理解任务中取得显著进展,但对人本场景的理解能力仍缺乏系统评估,主要受限于缺乏兼顾细粒度人体感知与高阶因果推理能力的综合基准。为此,本文提出Human-MME,一个精心设计的基准,用于更全面地评估多模态大模型在人本场景理解中的表现。相较于现有基准,其三大特点为:1)场景多样性,涵盖4个主视觉领域、15个二级领域和43个子领域,实现广泛场景覆盖;2)渐进式多维评估,从人体细粒度感知到高阶推理,包含8个维度、19,945对真实世界图像问答对及完整评估套件;3)高质量标注,构建自动化标注流程与人工标注平台,支持严格手动标注,确保评估精准可靠。该基准通过选择题、简答题、定位、排序、判断等题型及其组合,拓展了从单目标理解到多人多图互相关联理解的能力评估。对17个先进MLLMs的广泛实验有效揭示其局限性,为未来人本图像理解研究指明方向。所有数据与代码已开源:https://github.com/Yuan-Hou/Human-MME。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of comprehensive evaluation benchmarks that take into account both the human-oriented granular level and higher-dimensional causal reasoning ability. Such high-quality evaluation benchmarks face tough obstacles, given the physical complexity of the human body and the difficulty of annotating granular structures. In this paper, we propose Human-MME, a curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric scene understanding. Compared with other existing benchmarks, our work provides three key features: 1. Diversity in human scene, spanning 4 primary visual domains with 15 secondary domains and 43 sub-fields to ensure broad scenario coverage. 2. Progressive and diverse evaluation dimensions, evaluating the human-based activities progressively from the human-oriented granular perception to the higher-dimensional reasoning, consisting of eight dimensions with 19,945 real-world image question pairs and an evaluation suite. 3. High-quality annotations with rich data paradigms, constructing the automated annotation pipeline and human-annotation platform, supporting rigorous manual labeling to facilitate precise and reliable model assessment. Our benchmark extends the single-target understanding to the multi-person and multi-image mutual understanding by constructing the choice, short-answer, grounding, ranking and judgment question components, and complex questions of their combination. The extensive experiments on 17 state-of-the-art MLLMs effectively expose the limitations and guide future MLLMs research toward better human-centric image understanding. All data and code are available at https://github.com/Yuan-Hou/Human-MME.

多模态人本理解评估基准图像问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。