让机器像教练一样分析动作并给出详细改进建议。
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- 用多模态思维链构建完整推理过程,从识别动作到提出解决方案。
- 在动作分类和质量评估上分别提升2.7%和2.1%准确率,解释生成能力提高16.0%。
- 适合体育教学、智能健身设备等需要可解释反馈的场景。
在真实场景中评估人体动作是否规范并提供合理反馈以提升动作标准化至关重要但极具挑战。现有视频理解方法多关注动作的‘是什么’和‘在哪里’,难以满足需求;且多数数据集缺乏动作标准化程度标注,质量评估数据集也缺乏可解释性与具体建议。为此,我们定义了新的动作形式评估(AFA)任务,构建了大规模的健身与武术视频数据集CoT-AFA,包含多层级标注,支持全面视频分析。我们引入新颖的思维链(Chain-of-Thought)解释范式,不再仅提供孤立反馈,而是呈现从识别动作步骤、分析其效果到提出具体改进方案的完整推理过程。同时提出可解释健身评估框架Explainable Fitness Assessor,采用双路并行处理与动态门控机制融合视觉与语义信息,增强分析能力。实验表明,该方法在解释生成(如CIDEr提升16.0%)、动作分类(准确率+2.7%)和质量评估(准确率+2.1%)方面均有显著提升,展现了CoT-AFA在后续研究中的巨大潜力。数据集与代码已开源。
原文摘要 · Abstract (English)
Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. However, current video understanding methods are mainly concerned with what and where the action is, which is unable to meet the requirements. Meanwhile, most of the existing datasets lack the labels indicating the degree of action standardization, and the action quality assessment datasets lack explainability and detailed feedback. Therefore, we define a new Human Action Form Assessment (AFA) task, and introduce a new diverse dataset CoT-AFA, which contains a large scale of fitness and martial arts videos with multi-level annotations for comprehensive video analysis. We enrich the CoT-AFA dataset with a novel Chain-of-Thought explanation paradigm. Instead of offering isolated feedback, our explanations provide a complete reasoning process--from identifying an action step to analyzing its outcome and proposing a concrete solution. Furthermore, we propose a framework named Explainable Fitness Assessor, which can not only judge an action but also explain why and provide a solution. This framework employs two parallel processing streams and a dynamic gating mechanism to fuse visual and semantic information, thereby boosting its analytical capabilities. The experimental results demonstrate that our method has achieved improvements in explanation generation (e.g., +16.0% in CIDEr), action classification (+2.7% in accuracy) and quality assessment (+2.1% in accuracy), revealing great potential of CoT-AFA for future studies. Our dataset and source code is available at https://github.com/MICLAB-BUPT/EFA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。