构建3D足球场景理解新数据集,支持单目球体定位与场景重建。
SoccerNet-v3D: Leveraging Sports Broadcast Replays for 3D Scene Understanding
- 基于场线校准与多视角同步,实现足球场3D空间定位。
- 提出单图3D球体定位基线方法,精度达72.5% (mAP@10)。
- 适合体育视觉、3D场景重建方向研究者使用。
体育视频分析是计算机视觉的关键领域,可通过多视角对应关系实现精细的空间理解。本文提出SoccerNet-v3D和ISSIA-3D两个增强且可扩展的数据集,用于足球广播分析中的3D场景理解。它们在SoccerNet-v3和ISSIA基础上引入基于场线的相机标定与多视角同步,支持通过三角测量实现3D目标定位。我们设计了单目3D球体定位任务,基于真值2D球体标注进行三角化,并提供多种标定与重投影评估指标以检验标注质量。此外,提出一种单图像3D球体定位基线方法,利用相机标定与球体尺寸先验从单视角估计球体位置。为优化2D标注,还引入边界框优化技术,确保其与3D场景表示对齐。所提数据集建立3D足球场景理解新基准,提升体育分析中的时空建模能力。最后,公开代码以支持标注获取与数据集生成流程。
原文摘要 · Abstract (English)
Sports video analysis is a key domain in computer vision, enabling detailed spatial understanding through multi-view correspondences. In this work, we introduce SoccerNet-v3D and ISSIA-3D, two enhanced and scalable datasets designed for 3D scene understanding in soccer broadcast analysis. These datasets extend SoccerNet-v3 and ISSIA by incorporating field-line-based camera calibration and multi-view synchronization, enabling 3D object localization through triangulation. We propose a monocular 3D ball localization task built upon the triangulation of ground-truth 2D ball annotations, along with several calibration and reprojection metrics to assess annotation quality on demand. Additionally, we present a single-image 3D ball localization method as a baseline, leveraging camera calibration and ball size priors to estimate the ball's position from a monocular viewpoint. To further refine 2D annotations, we introduce a bounding box optimization technique that ensures alignment with the 3D scene representation. Our proposed datasets establish new benchmarks for 3D soccer scene understanding, enhancing both spatial and temporal analysis in sports analytics. Finally, we provide code to facilitate access to our annotations and the generation pipelines for the datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。