arXiv:2511.17045cs.CVcs.AI2025-11AAAI被引 4

构建首个球拍与球联合分析的数据集,推动网球等运动的视觉理解

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

  • 同时标注球和球拍的精细姿态,支持多模态融合分析
  • 引入交叉注意力机制,使球拍信息显著提升轨迹预测精度
  • 适用于运动分析、动态追踪及多模态建模的研究者

我们提出 RacketVision,一个面向乒乓球、网球和羽毛球的新型数据集与基准,用于推动体育视觉分析的发展。该数据集首次提供大规模、细粒度的球拍姿态标注,结合传统球体位置信息,支持对复杂人-物交互关系的研究。任务涵盖细粒度球体追踪、可动球拍姿态估计以及球体轨迹预测。对主流基线的评估揭示了一个关键发现:简单拼接球拍特征会降低性能,而采用交叉注意力机制才能充分挖掘其价值,实现超越强单模态基线的轨迹预测结果。RacketVision 为动态物体追踪、条件运动预测及体育领域的多模态分析提供了通用资源和有力起点。

原文摘要 · Abstract (English)

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a CrossAttention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multimodal analysis in sports. Project page at https://github.com/OrcustD/RacketVision

体育视觉多模态轨迹预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。