arXiv:2608.08736cs.AI2026-08被引 1

构建首个多模态大模型健身动作质量评估基准,揭示模型在识别错误上的瓶颈。

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

论文配图:FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
图 1 · 摘自论文原文
  • 设计统一的38类动作错误分类体系,覆盖6个质量维度。
  • 在2219段视频上测试,发现模型对错误定位准确率仍很低。
  • 适合研究多模态理解、智能健身或动作分析的学者使用。

健身动作质量评估(AQA)对智能训练至关重要,但多模态大语言模型(MLLMs)在此任务中的能力尚未充分探索。现有基准依赖特定动作标注且仅关注最终评分,难以揭示模型评估机制。我们提出FitAQA,一个系统性基准,包含2,219个视频和5,512个问答实例,覆盖30种自重动作。联合运动科学专家建立统一的形式错误分类体系,涵盖6个互补质量维度:对齐、对称、稳定、协调、节奏与完整性,共定义38类常见错误。该体系实现跨动作的统一评估框架。FitAQA进一步设定三项任务:感知(识别视觉证据)、判断(结合领域知识评估执行正确性)、时序定位(定位错误发生时间)。大量实验表明,当前MLLMs仍无法全面评估动作质量,也难以精准定位错误。受控实验显示,视觉感知是主要瓶颈——当提供真实感知证据后,判断性能显著提升。数据集与评估代码将公开发布。

原文摘要 · Abstract (English)

Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored. Existing benchmarks rely on action-specific annotation schemes and focus primarily on final assessment outputs, offering limited insight into how models assess exercise quality. We introduce FitAQA, a systematic benchmark for evaluating MLLMs in fitness AQA, containing 2,219 videos and 5,512 QA instances across 30 bodyweight exercises. In collaboration with experts in sports science, we develop a unified form error taxonomy that defines 38 recurring form errors within six complementary quality dimensions: alignment, symmetry, stability, coordination, tempo, and completeness. This taxonomy provides a shared assessment framework across different exercises. FitAQA further formulates three evaluation tasks: perception for recognizing relevant visual evidence, judgement for combining that evidence with domain knowledge to assess execution correctness, and temporal grounding for localizing form errors over time. Extensive evaluation shows that current MLLMs still struggle to assess exercise quality comprehensively and localize form errors precisely. Controlled experiments further indicate that visual perception is a key bottleneck, as judgement performance improves substantially when ground-truth perceptual evidence is provided. The dataset and evaluation code will be made publicly available.

动作评估多模态大模型健身

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。