用视觉语言模型构建新基准,精准评估动作理解能力
NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

- 三类任务+三重语义轴+三级难度,系统化评测动作理解
- 12个VLM模型测试发现传统评估无法暴露的深层缺陷
- 模型在细节判断上严重失准,适合研究者验证评估边界
可靠评估人类动作理解是推动具身AI、机器人和动画发展的基础。现有基准存在语义粒度粗、难度分布不均、标注质量有限和答案模糊等问题,难以诊断模型失败原因。为此,我们提出NextMotionQA,一个利用视觉语言模型(VLMs)实现半自动、专家验证的数据集。该基准包含三项互补任务:多选问答、视频描述和细粒度错误修正,每项任务沿三个核心语义轴设计,并分三级复杂度。对12个代表性VLM的全面评估揭示了传统单任务评估下隐匿的关键能力缺口。同时,近期研究尝试用VLM作为文本到动作生成的评价裁判;我们检验其在更难任务下的表现,发现其在粗粒度标准上与专家评分高度一致(Cohen's κ=0.70),但在细粒度部件级判断上严重失效(κ=0.10),验证了该范式在强条件下的有效性,也明确了其局限性。
原文摘要 · Abstract (English)
Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffer from coarse semantic granularity, undifferentiated difficulty, limited annotation quality, and pervasive answer ambiguity, leaving them unable to diagnose where current models fail. To bridge this gap, we introduce NextMotionQA, a comprehensive benchmark that leverages vision-language models (VLMs) for semi-automated, expert-verified dataset. NextMotionQA features three complementary tasks: multiple-choice question answering, video captioning, and fine-grained error correction. Each task is systematically structured across three core semantic axes and stratified into three task complexity levels. Our extensive evaluation of twelve representative VLMs uncovers critical capability gaps and weakness that remain invisible under conventional, single-task evaluations. In a complementary direction, recent work has begun using VLMs as judges for text-to-motion evaluation; we ask whether they show the same degradation under harder tasks. We find that VLMs align strongly with expert ratings on coarse criteria (Cohen's κ=0.70) but break down on fine-grained, part-level judgment (κ=0.10), validating the paradigm in its strong regime while clarifying its limits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。