arXiv:2606.20542cs.CV2026-06被引 1

构建首个大规模网球动作多视角数据集,用于评估单目视频转3D姿态的性能。

CalTennis: Large Multi-View Tennis Video Dataset and Benchmark of Monocular-to-3D Pose Estimation

论文配图:CalTennis: Large Multi-View Tennis Video Dataset and Benchmark of Monocular-to-3D Pose Estimation
图 1 · 摘自论文原文
  • 采用2-6个同步摄像头采集40名运动员51小时视频,每秒60帧
  • 现有模型在深度和脚部接触估计上表现不佳,但关节角度恢复较准
  • 提出脚法与稳定性新指标,揭示姿态估计未被关注的缺陷

Caltech Tennis Dataset(CalTennis)是首个面向真实场景下单目视频转3D姿态估计的大规模视频基准。该数据集包含超过1100万帧(51小时)的网球训练与比赛视频,来自40名运动员,由2至6台同步摄像头以60 Hz采样录制。其规模为现有真实场景人体运动数据集的10倍,为现有带动捕真值数据集的3倍,且是首个提供专家级运动多视角同步记录的大型基准。多视角设置支持低成本、免标注的单目3D姿态估计算法评估。我们提出了一种无需专业设备或技能的标准化采集协议,并实现全自动视频校准与同步。在CalTennis上对前沿单目3D姿态方法进行基准测试发现,尽管3D关节角度恢复已较为准确,但所有模型在深度估计与脚部接触判断上仍不稳定。我们进一步提出脚法与稳定性两个新评价指标,并定性分析身体形状不一致性。这些指标揭示了此前未被充分关注的失败模式,为姿态估计与动作分析指明了具体改进方向。

原文摘要 · Abstract (English)

The Caltech Tennis Dataset (CalTennis) is a large-scale video benchmark for evaluating monocular-to-3D pose estimation in the wild. CalTennis comprises over 11 million frames (51 hours) of tennis practice and match play from 40 players, captured with 2-6 synchronized cameras at 60 Hz. It is 10 times larger than existing in-the-wild human motion video datasets and 3 times larger than existing MOCAP-ground-truthed datasets, and it is the first large-scale benchmark to provide synchronized multi-view recordings of expert athletic motion. The multi-view setup enables inexpensive, label-free evaluation of monocular-to-3D pose estimation algorithms. We describe a simple, standardized protocol that enables data collection without specialized equipment or expertise, along with fully automated video calibration and synchronization. Benchmarking state-of-the-art monocular-to-3D pose methods on CalTennis, we find that while 3D joint angle recovery is now quite accurate, all models struggle to estimate depth and foot contact consistently. We further propose two novel performance metrics, footwork and stability, as well as qualitatively study body shape inconsistency. These metrics expose previously underexplored failure modes and point to concrete opportunities for improvement in pose estimation and action analysis.

动作识别3D姿态估计多视角数据集网球分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。