评测视觉语言模型在动作质量评估中的表现,发现其效果仅略高于随机。
Can Vision Language Models Judge Action Quality? An Empirical Evaluation

- 测试多种视觉语言模型在健身、花样滑冰等领域的动作评估表现
- 顶级模型如Gemini 3.1 Pro仅略高于随机准确率(约50%)
- 模型存在视觉证据忽视和语言框架敏感等系统性偏差
动作质量评估(AQA)在物理治疗、体育训练和竞技评判中具有广泛应用。尽管视觉语言模型(VLMs)在该领域前景广阔,但其实际表现仍不明确。本文对最先进的VLMs在多个活动领域(如健身、花样滑冰、跳水)、任务类型、表示方式和提示策略下进行了全面评估。基线结果显示,Gemini 3.1 Pro、Qwen3-VL和InternVL3.5模型的性能仅略高于随机猜测(约50%)。虽然引入骨骼信息、指令定位、推理结构和上下文学习等策略带来局部提升,但均未表现出持续有效性。对预测分布的分析揭示两种系统性偏差:倾向于无视视觉证据而预测正确执行,以及对语言表达形式高度敏感。通过对比式任务重构以缓解偏差,改善有限,表明模型在精细动作质量判断上的局限远超上述偏差,指向根本性挑战。研究建立了未来VLM-based AQA研究的严格基准,并为实际部署前需解决的失败模式提供了可操作的路线图。
原文摘要 · Abstract (English)
Action Quality Assessment (AQA) has broad applications in physical therapy, sports coaching, and competitive judging. Although Vision Language Models (VLMs) hold considerable promise for AQA, their actual performance in this domain remains largely uncharacterised. We present a comprehensive evaluation of state-of-the-art VLMs across activity domains (e.g. fitness, figure skating, diving), tasks, representations, and prompting strategies. Baseline results reveal that Gemini 3.1 Pro, Qwen3-VL and InternVL3.5 models perform only marginally above random chance, and although strategies such as incorporation of skeleton information, grounding instructions, reasoning structures and in-context learning lead to isolated gains, none is consistently effective. Analysis of prediction distributions uncovers two systematic biases: a tendency to predict correct execution regardless of visual evidence, and a sensitivity to superficial linguistic framing. Reformulating tasks contrastively to mitigate these biases yields minimal improvement, suggesting that the models' limitations go beyond these biases, pointing to a fundamental difficulty with fine-grained movement quality assessment. Our findings establish a rigorous baseline for future VLM-based AQA research and provide an actionable outline for failure modes requiring mitigation prior to reliable real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。