arXiv:2507.04909cs.CVcs.AI2025-07被引 3

构建首个全面评估大模型理解人类视频能力的基准

HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding

  • 设计13项任务覆盖感知到认知的多维度评测
  • 支持从10秒到30分钟视频的时序理解能力测试
  • 适合研究人机交互、视频理解与大模型评估的学者

多模态大语言模型在图像和视频理解任务中取得显著进展,但其对以人为中心的视频数据的理解能力仍缺乏深入探索,主要受限于高质量评估基准的缺失。现有基准多聚焦视频生成质量与动作识别,忽视了人类场景中必要的感知与认知能力,且常采用单题问答模式和简单评价指标。为此,我们提出HV-MMBench,一个精心构建的现代人类中心视频理解基准。相比现有基准,本工作具有四大特点:(1) 多维度评估:涵盖13项任务,从基础属性感知(如年龄估计、情绪识别)到高级认知推理(如社会关系预测、意图预测),实现模型能力的全面评估;(2) 多样化数据类型:包含选择题、填空题、判断题和开放式问题,结合多种评价指标,更准确稳健地反映模型表现;(3) 多领域视频覆盖:涵盖50个不同视觉场景,支持细粒度场景差异下的综合评估;(4) 跨时长覆盖:视频时长从短时(10秒)到长时(最高30分钟),支持对模型时序推理能力的系统分析。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks involving both images and videos. However, their capacity to comprehend human-centric video data remains underexplored, primarily due to the absence of comprehensive and high-quality evaluation benchmarks. Existing human-centric benchmarks predominantly emphasize video generation quality and action recognition, while overlooking essential perceptual and cognitive abilities required in human-centered scenarios. Furthermore, they are often limited by single-question paradigms and overly simplistic evaluation metrics. To address above limitations, we propose a modern HV-MMBench, a rigorously curated benchmark designed to provide a more holistic evaluation of MLLMs in human-centric video understanding. Compared to existing human-centric video benchmarks, our work offers the following key features: (1) Diverse evaluation dimensions: HV-MMBench encompasses 13 tasks, ranging from basic attribute perception (e.g., age estimation, emotion recognition) to advanced cognitive reasoning (e.g., social relationship prediction, intention prediction), enabling comprehensive assessment of model capabilities; (2) Varied data types: The benchmark includes multiple-choice, fill-in-blank, true/false, and open-ended question formats, combined with diverse evaluation metrics, to more accurately and robustly reflect model performance; (3) Multi-domain video coverage: The benchmark spans 50 distinct visual scenarios, enabling comprehensive evaluation across fine-grained scene variations; (4) Temporal coverage: The benchmark covers videos from short-term (10 seconds) to long-term (up to 30min) durations, supporting systematic analysis of models temporal reasoning abilities across diverse contextual lengths.

视频理解多模态大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。