arXiv:2606.27999cs.CV2026-06

首个评估视频中人体运动全局轨迹与朝向推理的基准,填补了现有模型在复杂动作理解上的空白。

HumanMoveVQA: Can Video MLLMs reason about human movement in videos?

论文配图:HumanMoveVQA: Can Video MLLMs reason about human movement in videos?
图 1 · 摘自论文原文
  • 构建基于第一帧坐标系的3D运动追踪系统,实现世界一致的轨迹建模
  • 生成超1万条结构化问答对,覆盖7类运动推理任务,验证主流模型存在显著能力缺口
  • 提出可微调的开放基线,证明运动理解可通过特定监督有效提升,适合视频理解研究者

尽管多模态大语言模型在高层视频理解方面取得快速进展,但一个根本瓶颈依然存在:这些模型将复杂的肢体运动简化为粗粒度语义标签。现有基准大多关注场景级事件或局部关节运动,未能深入探查人体在时空中的全局运动(如轨迹与朝向变化)。我们提出HumanMoveVQA,首个专为评估外视角下全局轨迹与朝向推理设计的综合性基准。该基准采用以第一帧为锚点的世界坐标系,保持相对于固定起点的平移与旋转一致性。我们设计了一套可扩展的多阶段流程,将2D视频观测提升为世界一致的3D运动轨迹,并生成超过10,000条结构化问答对,涵盖运动聚合、序列排序、轨迹级推断等七类推理任务。大规模评估揭示当前顶级专有模型在深层人体运动理解上存在显著能力差距。然而我们证明,这一问题具有可学习性:通过使用我们的目标导向、世界一致的监督信号对开源基线进行微调,实现了显著性能提升。HumanMoveVQA为下一代具身感知视频理解模型奠定了严格的几何基础。

原文摘要 · Abstract (English)

Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motion into coarse semantic labels. Existing benchmarks mostly focus on scene-centric events or local joint articulations, failing to probe global human motion in space over time (trajectory and orientation changes). We introduce HumanMoveVQA, the first comprehensive benchmark designed to evaluate global trajectory and orientation reasoning from an exocentric perspective. Our benchmark utilizes a first-frame anchored world coordinate system, preserving translation and rotation relative to a fixed starting point. We propose a scalable, multi-stage pipeline that lifts 2D video observations into world-consistent 3D motion tracks to generate over 10K structured question-answer pairs across seven reasoning categories, including motion aggregation, sequential ordering, and trajectory-level inference. Our extensive evaluation reveals a critical capability gap in state-of-the-art proprietary models on deep human motion understanding. However, we demonstrate that this is a learnable problem; by fine-tuning an open-source baseline with our targeted, world-consistent supervision, we achieve a significant improvement. HumanMoveVQA establishes a rigorous geometric foundation for developing next-generation, movement-aware video understanding models.

视频理解运动推理多模态3D轨迹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。