跨运动通用技能评估,生成可操作反馈
Learning Skill-Attributes for Transferable Assessment in Video
- 从视频中自动识别通用技能属性(如平衡、手位)
- 跨运动场景提升评估准确率最高达60%相对提升
- 适合需要低成本高质量反馈的体育训练场景
从视频中进行技能评估需判断人体动作质量并给出改进建议。当前模型多针对单一运动,受限于长尾运动中专家标注的稀缺与高成本。为此,本文提出CrossTrainer方法,自动发现跨运动通用的技能属性(如平衡、控制、手部位置),并训练多模态语言模型,为新视频生成可操作建议(如“抬手更高以增加力量”)及熟练度等级(如“初级专家”)。在多个数据集上验证,该方法在跨运动(迁移)和同运动(域内)设置中均取得最高60%相对性能提升。通过抽象共性行为特征,所提视频表征显著优于现有技术,增强了多模态大模型的应用能力。
原文摘要 · Abstract (English)
Skill assessment from video entails rating the quality of a person's physical performance and explaining what could be done better. Today's models specialize for an individual sport, and suffer from the high cost and scarcity of expert-level supervision across the long tail of sports. Towards closing that gap, we explore transferable video representations for skill assessment. Our CrossTrainer approach discovers skill-attributes, such as balance, control, and hand positioning -- whose meaning transcends the boundaries of any given sport, then trains a multimodal language model to generate actionable feedback for a novel video, e.g., "lift hands more to generate more power" as well as its proficiency level, e.g., early expert. We validate the new model on multiple datasets for both cross-sport (transfer) and intra-sport (in-domain) settings, where it achieves gains up to 60% relative to the state of the art. By abstracting out the shared behaviors indicative of human skill, the proposed video representation generalizes substantially better than an array of existing techniques, enriching today's multimodal large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。