arXiv:2605.03485cs.CVcs.AI2026-05

构建多维人像感知与推理评测基准,提升大模型对人类行为的理解能力。

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

论文配图:MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models
图 1 · 摘自论文原文
  • 设计多层级数据集与自动化标注流水线,实现高质量人像属性标注。
  • 使用强化学习优化困难样本,显著提升模型在复杂场景下的推理表现。
  • 适合研究视觉语言模型中人类行为理解的学者与开发者使用。

真实世界应用如影视分析与虚拟数字人需要对人类进行多维度理解,但现有视觉语言模型(LVLM)评测多聚焦单一任务,缺乏细粒度、以人为本的评估。本文提出MHPR,一个覆盖个体、多人及人-物交互的联合感知-推理基准。其包含四类数据:带注释原始数据(C-RD)、监督微调数据(SFT-D)、强化学习数据(RL-D)和测试数据(T-D),并配备自动化标题/问答生成流水线(ACVG),通过类别属性分解、特定属性重写与多模型投票保障标注质量与可扩展性。我们在细粒度属性(外貌、着装、姿态、部位)与高层语义(社交关系、动作语义、空间关系、意图与功能)上评估主流模型。结果表明:1)格式对齐的SFT数据显著提升指令遵循与稳定性;2)基于错误案例分析构建的挑战性RL数据进一步增强模型在困难实例上的感知与推理能力;3)用MHPR训练Qwen2.5-VL-7B模型后,性能接近更大模型。我们开源了ACVG与MHPR,以推动可复现、可扩展的人类中心感知与推理研究。

原文摘要 · Abstract (English)

Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchmarks largely focus on single-task settings and lack fine-grained, human-centric evaluation. In this work, we introduce MHPR, a comprehensive benchmark for joint perception-reasoning over human-centric scenes spanning individual, multi-person, and human-object interaction dimensions. MHPR comprises a multi-level data design-Captioned Raw Data (C-RD), Supervised Fine-Tuning Data (SFT-D), Reinforcement Learning Data (RL-D), and Test Data (T-D)-together with an automated caption/VQA generation pipeline (ACVG) that performs category-wise attribute decomposition, attribute-specific rewriting, and multi-model voting to ensure high-quality, scalable annotations. We evaluate state-of-the-art vision-language models on fine-grained attributes (appearance, clothing, pose, parts) and high-level semantics (social relations, action semantics, spatial relations, intent and functionality). Our findings show that: 1) format-aligned SFT data substantially improves instruction following and stability; 2) challenge-focused RL data derived from bad-case analysis further enhances perception and reasoning on difficult instances; and 3) training Qwen2.5-VL-7B with MHPR yields significant gains, achieving near-parity with considerably larger models. We release ACVG and MHPR to facilitate reproducible, extensible research on human-centric perception and reasoning.

视觉语言模型人像理解评测基准强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。