构建人类中心理解新基准,提升多模态大模型对复杂场景的感知能力
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
- 设计人类中心理解评测基准HERM-Bench,评估多模态模型在复杂场景下的表现
- 推出包含10万条多粒度标注的HERM-100K数据集,增强模型训练
- 提出HERM-7B模型,在多项人类中心任务上超越现有模型
多模态大语言模型(MLLMs)在视觉理解和指令遵循方面取得显著进展,为广泛的人类中心应用场景提供了可能。然而,现有图像-文本数据难以支持多粒度信息的精确模态对齐与融合,这制约了人类中心视觉理解的发展。本文提出HERM-Bench,用于评估MLLMs在人类中心理解方面的性能。研究揭示现有模型在复杂人类中心场景中存在明显局限。为此,我们构建了包含10万条样本、具有多层次人类中心标注的HERM-100K数据集,并基于该数据训练出HERM-7B模型。在HERM-Bench上的评估显示,HERM-7B在多个维度上显著优于现有MLLMs,反映出当前训练数据在人类中心视觉理解上的不足。本研究强调专用数据集与评测基准对推动人类中心理解能力的重要性。
原文摘要 · Abstract (English)
The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centric scenarios. However, existing image-text data may not support the precise modality alignment and integration of multi-grained information, which is crucial for human-centric visual understanding. In this paper, we introduce HERM-Bench, a benchmark for evaluating the human-centric understanding capabilities of MLLMs. Our work reveals the limitations of existing MLLMs in understanding complex human-centric scenarios. To address these challenges, we present HERM-100K, a comprehensive dataset with multi-level human-centric annotations, aimed at enhancing MLLMs' training. Furthermore, we develop HERM-7B, a MLLM that leverages enhanced training data from HERM-100K. Evaluations on HERM-Bench demonstrate that HERM-7B significantly outperforms existing MLLMs across various human-centric dimensions, reflecting the current inadequacy of data annotations used in MLLM training for human-centric visual understanding. This research emphasizes the importance of specialized datasets and benchmarks in advancing the MLLMs' capabilities for human-centric understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。