arXiv:2603.25533cs.CV2026-03中稿 · ed

首个全场比赛密集标注的羽毛球数据集,支持完整赛事分析与精准击球描述。

BFMD: A Full-Match Badminton Dense Dataset for Dense Shot Captioning

  • 构建包含19场全比赛、超20小时视频的密集标注数据集
  • 涵盖16,751次击球事件,每条都有击球类型与动作轨迹标注
  • 提出语义反馈机制,提升多模态击球描述的准确性与连贯性

理解羽毛球战术动态需分析完整比赛而非孤立片段。现有数据集多聚焦短片段或特定任务标注,缺乏全场比赛的密集多模态标注,难以生成准确击球描述或进行比赛级分析。为此,我们推出首个羽毛球全场比赛密集标注(BFMD)数据集,包含19场广播比赛(单打与双打),覆盖超过20小时比赛内容,包含1,687个回合与16,751次击球事件,每条击球均有击球描述标注。数据提供层级化标注:比赛段落、回合事件及回合级多模态信息,包括击球类型、球体轨迹、球员姿态关键点与击球描述。我们设计基于VideoMAE的多模态描述框架,引入语义反馈机制,利用击球语义引导生成过程,提升语义一致性。实验表明,多模态建模与语义反馈显著优于仅用RGB图像的基线模型。进一步通过分析全场比赛中战术模式的时序演变,展示了该数据集在战术研究中的潜力。

原文摘要 · Abstract (English)

Understanding tactical dynamics in badminton requires analyzing entire matches rather than isolated clips. However, existing badminton datasets mainly focus on short clips or task-specific annotations and rarely provide full-match data with dense multimodal annotations. This limitation makes it difficult to generate accurate shot captions and perform match-level analysis. To address this limitation, we introduce the first Badminton Full Match Dense (BFMD) dataset, with 19 broadcast matches (including both singles and doubles) covering over 20 hours of play, comprising 1,687 rallies and 16,751 hit events, each annotated with a shot caption. The dataset provides hierarchical annotations including match segments, rally events, and dense rally-level multimodal annotations such as shot types, shuttle trajectories, player pose keypoints, and shot captions. We develop a VideoMAE-based multimodal captioning framework with a Semantic Feedback mechanism that leverages shot semantics to guide caption generation and improve semantic consistency. Experimental results demonstrate that multimodal modeling and semantic feedback improve shot caption quality over RGB-only baselines. We further showcase the potential of BFMD by analyzing the temporal evolution of tactical patterns across full matches.

羽毛球多模态密集标注视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。