arXiv:2603.25163cs.CV2026-03被引 1

用36万条运动视频训练模型,提升对动作细节的理解能力。

SportSkills: Physical Skill Learning from Sports Instructional Videos

  • 构建首个面向动作学习的运动教学视频数据集,含63万段视觉演示。
  • 相同模型在新数据上性能最高提升4倍,能区分细微动作差异。
  • 可实现错误动作匹配教学视频,适合教练与健身用户个性化指导。

现有大规模视频数据集聚焦通用人类活动,缺乏精细动作的深度覆盖。我们提出SportSkills,首个面向物理技能学习的大规模野外运动教学视频数据集。该数据集包含超过36万条教学视频,涵盖55种不同运动,共63万段视觉示范,并配有解说文字说明动作背后的技巧。通过一系列实验,我们证明SportSkills使模型能够理解精细动作间的差异,相同模型在传统活动中心数据集上训练时,性能最高提升4倍。更重要的是,基于SportSkills,我们首次提出错误条件下的教学视频检索任务,连接表征学习与可操作反馈生成(例如:“我执行了一个动作,该看哪个视频来改进?”)。专业教练的正式评估显示,该方法显著提升了视频模型为用户查询提供个性化视觉指导的能力。

原文摘要 · Abstract (English)

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared towards physical skill learning with in-the-wild video. SportSkills has more than 360k instructional videos containing more than 630k visual demonstrations paired with instructional narrations explaining the know-how behind the actions from 55 varied sports. Through a suite of experiments, we show that SportSkills unlocks the ability to understand fine-grained differences between physical actions. Our representation achieves gains of up to 4x with the same model trained on traditional activity-centric datasets. Crucially, building on SportSkills, we introduce the first large-scale task formulation of mistake-conditioned instructional video retrieval, bridging representation learning and actionable feedback generation (e.g., "here's my execution of a skill; which video clip should I watch to improve it?"). Formal evaluations by professional coaches show our retrieval approach significantly advances the ability of video models to personalize visual instructions for a user query.

动作识别视频理解个性化学习运动科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。