用自动生成的测试题评估大模型理解人类情感行为的能力
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
- 用自动化流程生成高质量视频标注和难题,减少人工成本
- 30个主流多模态大模型在16项任务中均未达人类水平,尤其情绪识别差
- 适合研究视频理解、社会智能模型的学者和开发者
评估多模态大语言模型(MLLMs)在人类中心视频理解方面的能力仍具挑战,现有基准常忽略情感、行为及跨模态对齐的细微之处。我们提出HumanVBench,一个涵盖16个细粒度任务的综合性视频评测基准。其核心是新颖且可扩展的自动构建方法,包含两条自动化流水线:分别合成高质量视频标注与具有挑战性的多项选择题,仅需极少人工参与。通过利用先进模型进行标注,并将模型错误系统性转化为合理干扰项,该框架可通用化生成精细评估集。我们在HumanVBench上对30个领先MLLMs进行广泛评估,发现其在感知细微情绪及对齐语音与视觉线索方面存在显著缺陷,即使顶级专有模型也未能达到人类表现。我们已开源HumanVBench及合成流水线,以推动更具备社会智能的视频MLLM发展。
原文摘要 · Abstract (English)
Evaluating the nuanced human-centric video understanding capabilities of Multimodal Large Language Models (MLLMs) remains a great challenge, as existing benchmarks often overlook the intricacies of emotion, behavior, and cross-modal alignment. We introduce HumanVBench, a comprehensive video benchmark designed to rigorously probe these capabilities across 16 fine-grained tasks. A cornerstone of our work is a novel and scalable benchmark construction methodology, featuring two automated pipelines that synthesize high-quality video annotations and challenging multiple-choice questions with minimal human labor. By leveraging state-of-the-art models for annotation and systematically converting model-induced errors into plausible distractors, our framework provides a generalizable ``machine'' for creating nuanced evaluation suites. Our extensive evaluation of 30 leading MLLMs on HumanVBench reveals critical deficiencies, particularly in perceiving subtle emotions and aligning speech with visual cues, with even top proprietary models falling short of human performance. We open-source HumanVBench and our synthesis pipelines to catalyze the development of more socially intelligent and capable video MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。