构建首个单目视频表征乒乓球4D动态的大型数据集与重建流程。
TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos

- 先3D还原球轨迹再分割时间片段,突破遮挡与视角变化限制。
- 140+小时高质量数据含球速、旋转、人体姿态等多模态标注。
- 适用于虚拟回放、球员分析及机器人学习,性能远超现有方法。
我们提出TT4D,一个大规模、高保真度的乒乓球4D重建数据集。该数据集基于单目转播视频,包含超过140小时的单打与双打比赛重建结果,涵盖多模态标注:高质量相机标定、精确的3D球位置、球旋转、时间分段及随时间变化的3D人体网格。丰富的数据为虚拟回放、深度球员分析和机器人学习提供了新基础。其规模与精度通过一种新颖的重建流水线实现:以往方法先根据2D球轨迹进行片段分割,再进行重建,但2D分割在遮挡和视角变化下失效。我们反向设计,先用学习的提升网络将完整未分割的2D球轨迹提升至3D,再据此可靠地进行时间分段。该提升网络还能推断球旋转、处理不可靠检测,在高遮挡情况下仍成功重建轨迹。该‘先提升后分割’设计使本流水线成为唯一可从通用视角单目转播视频中重建乒乓球比赛的方法。我们通过两个下游任务验证数据集保真度:准确估计击球瞬间球拍姿态与速度,以及训练生成高水平对打回合的生成模型。
原文摘要 · Abstract (English)
We present TT4D, a large-scale, high-fidelity table tennis dataset. It provides $140+$ hours of reconstructed singles and doubles gameplay from monocular broadcast videos, featuring multimodal annotations like high-quality camera calibrations, precise 3D ball positions, ball spin, time segmentation, and 3D human meshes over time. This rich data provides a new foundation for virtual replay, in-depth player analysis, and robot learning. The dataset's combination of scale and precision is achieved through a novel reconstruction pipeline. Prior methods first partition a game sequence into individual shot segments based on the 2D ball track, and only then attempt reconstruction. However, 2D-based time segmentation collapses under occlusion and varied camera viewpoints, preventing reliable reconstruction. We invert this paradigm by first lifting the entire unsegmented 2D ball track to 3D through a learned lifting network. This 3D trajectory then allows us to reliably perform time segmentation. The learned lifting network also infers the ball's spin, handles unreliable ball detections, and successfully reconstructs the ball trajectory in cases of high occlusion. This lift-first design is necessary, as our pipeline is the only method capable of reconstructing table tennis gameplay from general-view broadcast monocular videos. We demonstrate the dataset's fidelity through two downstream tasks: estimating the racket's pose \& velocity at impact, and training a generative model of competitive rallies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。