根据动作差异生成纠正指令,帮助用户改进运动表现。
CigTime: Corrective Instruction Generation Through Inverse Motion Editing
- 输入当前动作与目标动作,逆向生成指导文本
- 在多种运动场景中显著优于基线模型
- 适合体育教学与技能训练场景使用
近年来,将自然语言与人体动作关联的模型在基于文本指令的动作生成与编辑方面展现出巨大潜力。针对体育教练和运动技能学习的应用需求,我们研究了逆问题:基于动作编辑与生成模型,生成纠正性指导文本。本文提出一种新方法,给定用户的当前动作(源)和期望动作(目标),自动生成指导性文本以引导用户达成目标动作。利用大语言模型生成纠正文本,并借助现有动作生成与编辑框架构建三元组数据集(源动作、目标动作、纠正文本)。基于该数据,我们提出一种新型动作-语言模型用于生成纠正指令。在多样化应用场景中,通过定性和定量评估均表明,该方法显著优于基线模型,验证了其在教学指导中的有效性,可提供基于文本的精准动作纠正与性能提升建议。
原文摘要 · Abstract (English)
Recent advancements in models linking natural language with human motions have shown significant promise in motion generation and editing based on instructional text. Motivated by applications in sports coaching and motor skill learning, we investigate the inverse problem: generating corrective instructional text, leveraging motion editing and generation models. We introduce a novel approach that, given a user's current motion (source) and the desired motion (target), generates text instructions to guide the user towards achieving the target motion. We leverage large language models to generate corrective texts and utilize existing motion generation and editing frameworks to compile datasets of triplets (source motion, target motion, and corrective text). Using this data, we propose a new motion-language model for generating corrective instructions. We present both qualitative and quantitative results across a diverse range of applications that largely improve upon baselines. Our approach demonstrates its effectiveness in instructional scenarios, offering text-based guidance to correct and enhance user performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。