构建首个拳击动作识别与定位多模态基准数据集
BoxingVI: A Multi-Modal Benchmark for Boxing Action Recognition and Localization
- 从20段公开拳击视频中提取6915个高质量拳击片段
- 涵盖6类拳法,覆盖不同运动员、角度和动作风格
- 适合实时动作识别与自动训练分析研究
近年来,利用计算机视觉进行搏击运动的精准分析日益受到关注,但因动作动态性强、环境不固定,高质量数据集的构建仍是主要瓶颈。本文提出一个全面且标注精细的视频数据集,专用于拳击中的出拳检测与分类。数据集包含6915个高质拳击片段,来自20段公开的YouTube对练视频,涉及18名不同运动员,分为六类拳法。每个片段均人工分割并标注,确保时间边界精确、类别一致,涵盖丰富的运动风格、摄像角度和体型差异。该数据集旨在支持低资源、非约束环境下基于视觉的动作识别研究,通过提供多样化的拳击样本,推动拳击及类似领域在动作分析、自动教练和表现评估方面的进展。
原文摘要 · Abstract (English)
Accurate analysis of combat sports using computer vision has gained traction in recent years, yet the development of robust datasets remains a major bottleneck due to the dynamic, unstructured nature of actions and variations in recording environments. In this work, we present a comprehensive, well-annotated video dataset tailored for punch detection and classification in boxing. The dataset comprises 6,915 high-quality punch clips categorized into six distinct punch types, extracted from 20 publicly available YouTube sparring sessions and involving 18 different athletes. Each clip is manually segmented and labeled to ensure precise temporal boundaries and class consistency, capturing a wide range of motion styles, camera angles, and athlete physiques. This dataset is specifically curated to support research in real-time vision-based action recognition, especially in low-resource and unconstrained environments. By providing a rich benchmark with diverse punch examples, this contribution aims to accelerate progress in movement analysis, automated coaching, and performance assessment within boxing and related domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。