arXiv:2608.26094cs.CVcs.AI2026-08

用肌肉动态+动作分解,实现健身动作的精准评估与指导。

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

论文配图:MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
图 1 · 摘自论文原文
  • 融合视频、姿态、肌电等多模态数据,建模动作与肌肉活动关系。
  • 构建包含7500+样本的大型基准数据集,支持细粒度错误分析。
  • 适合健身、康复领域研究者,推动物理AI中可解释动作理解发展。

现有动作质量评估(AQA)数据集和方法主要依赖视觉输入(如RGB图像与姿态),忽视肌肉力学等生理动态,常将动作视为整体模式,难以提供精细化、生物力学基础的反馈。本文提出MyoMechanix,一个面向负重动作的多模态生态系统,实现运动与肌肉活动的对齐。该数据集由专家标注,包含38名受试者完成20种动作的7500+样本,同步采集多视角RGB视频、3D姿态、sEMG及其它生理信号,是目前最大的多模态AQA基准。我们进一步构建健身知识图谱(FKG),将专家标注结构化为动作、阶段、关键步骤、错误与纠正建议之间的关系,支持组合式评分与可解释评估。基于此,我们开发了CUBIST(组合本体推理引擎),通过分解-分析-重构实现细粒度错误归因与反馈生成。同时建立MyoMechanix-AQA、MyoMechanix-VideoQA及新型MyoMechanix-Video2EMG任务。实验表明,多模态感知与结构化表示显著提升性能、可解释性与错误归因能力,CUBIST达当前最优;VideoQA增强语言引导的动作理解;Video2EMG显示视频可替代昂贵的肌电传感。MyoMechanix推动技能活动理解向生物力学驱动、多模态、组合推理演进,适用于健身、康复、医疗与机器学习中的物理人工智能应用。

原文摘要 · Abstract (English)

Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic patterns. These limitations hinder fine-grained, biomechanically grounded feedback. We introduce MyoMechanix, a multimodal ecosystem for weight-loaded actions that aligns motion with muscle activity. Expert-annotated, it contains 7,500+ samples of 20 actions from 38 subjects, with synchronized multiview RGB video, 3D pose, sEMG, and additional physiological signals, forming the largest multimodal AQA benchmark to date. We further construct the Fitness Knowledge Graph (FKG), which organizes expert annotations into structured relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment. Building on these representations, we develop CUBIST (Compositional Ontological Reasoning Engine), which performs decomposition-analysis-recomposition for fine-grained error attribution and feedback generation. We also establish MyoMechanix-AQA, MyoMechanix-VideoQA, and a novel MyoMechanix-Video2EMG task. Experiments show that multimodal sensing and structured representations improve performance, interpretability, and error attribution, with CUBIST achieving state-of-the-art results; VideoQA enhances language-grounded action understanding; and Video2EMG suggests video-based alternatives to costly EMG sensing. MyoMechanix advances skilled activity understanding toward biomechanically grounded, multimodal, and compositional reasoning for Physical AI applications in fitness, rehabilitation, healthcare, and machine learning. Project page: https://haoyin116.github.io/MyoMechanix/

动作评估多模态肌电物理AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。