构建多层级人类场景评估体系,揭示大模型在理解人类行为上的短板。
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- 分感知、理解、推理三层次设计评测任务,覆盖9个维度的视觉认知能力
- 超6000道人工验证题测试模型在空间、时间与心理建模上的不足
- 提供带视觉证据链的思维过程标注,助力研究模型决策机制
为推动多模态大模型向人类水平智能迈进,我们提出HumanPCR,一个面向多样化人类相关视觉场景的评测体系。该体系涵盖感知(Human-P)、理解(Human-C)和推理(Human-R)三个层级:Human-P与Human-C包含超过6000道人工验证的多项选择题,覆盖9个维度的任务,包括现有基准常忽略的关键技能;Human-R则提供一项手动精心设计的视频推理测试,要求模型整合多源视觉证据、主动提取问题提示之外的上下文信息,并运用类人专业经验。每道题附有人工标注的思维链(CoT)及关键视觉证据,支持后续研究。对30余种顶尖模型的广泛评估显示,模型在细节空间感知、时间理解与心智建模任务中面临显著挑战。对Human-R的分析表明,模型难以从多样人类场景中主动提取关键视觉证据,且过度依赖查询引导的检索策略。即使采用扩大视觉上下文或测试时思维等先进方法,提升效果也有限。我们希望HumanPCR及其发现能推动多模态模型的开发、评估与以人为本的应用。
原文摘要 · Abstract (English)
The aspiration for artificial general intelligence, fueled by the rapid progress of multimodal models, demands human-comparable performance across diverse environments. We propose HumanPCR, an evaluation suite for probing MLLMs' capacity about human-related visual contexts across three hierarchical levels: Perception, Comprehension, and Reasoning (denoted by Human-P, Human-C, and Human-R, respectively). Human-P and Human-C feature over 6,000 human-verified multiple choice questions, assessing massive tasks of 9 dimensions, including but not limited to essential skills frequently overlooked by existing benchmarks. Human-R offers a challenging manually curated video reasoning test that requires integrating multiple visual evidences, proactively extracting context beyond question cues, and applying human-like expertise. Each question includes human-annotated Chain-of-Thought (CoT) rationales with key visual evidence to support further research. Extensive evaluations on over 30 state-of-the-art models exhibit significant challenges in human-centric visual understanding, particularly in tasks involving detailed space perception, temporal understanding, and mind modeling. Moreover, analysis of Human-R reveals the struggle of models in extracting essential proactive visual evidence from diverse human scenes and their faulty reliance on query-guided retrieval. Even with advanced techniques like scaling visual contexts and test-time thinking yield only limited benefits. We hope HumanPCR and our findings will advance the development, evaluation, and human-centric application of multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。