首个跨多视角3D点追踪基准,支持相机运动下的长期追踪。
TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

- 构建跨多视角、相机运动下的3D点追踪基准,含109,769条轨迹。
- 30余种基线方法均表现不佳,现有方法未优于单目追踪器。
- 揭示几何重建是3D点追踪的核心瓶颈,适合多视角感知研究者。
多摄像头系统在机器人、AR/VR和自动驾驶中日益实用,因其互补视角可降低深度模糊并提升遮挡下的可见性。然而,现有点追踪基准集中于单视频或静态多相机阵列,缺乏对相机运动下长期跨视图3D点追踪的评估。我们提出TAPVid-MV(Tracking Any Point in Video across Multiple Views),首个面向此场景的基准。其包含284个序列、1,142个校准摄像头流及109,769条跨七个子集(涵盖室内、室外、机器人、人体活动、驾驶及合成程序场景)的点轨迹。轨迹通过特定数据集辅助模态生成:传感器深度、LiDAR、SLAM与SfM点、人体网格、带姿态的对象网格及仿真。每条序列和轨迹均经人工视觉验证。在超过30种基线方法上,无一接近解决该任务。令人意外的是,现有多视角点追踪器并未持续优于单目追踪器。通过在同一数据集上联合评估重建与点追踪,TAPid-MV有助于区分几何恢复误差与点对应误差。联合分析表明,几何恢复是准确3D点追踪的主要瓶颈。除多视角3D点追踪外,释放的标注还支持单目2D/3D点追踪、未来轨迹预测与4D重建。
原文摘要 · Abstract (English)
Multi-camera systems are increasingly practical for robotics, AR/VR, and autonomous driving because complementary views reduce depth ambiguity and preserve visibility under occlusion. Existing point-tracking benchmarks, however, focus on a single video or static multi-camera rigs. None test long-term 3D point tracking across several synchronized views under camera motion. We introduce TAPVid-MV (Tracking Any Point in Video across Multiple Views), the first benchmark for this setting. It contains a curated set of 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks across seven subsets spanning indoor and outdoor domains, from robotics and human activity to driving and synthetic procedural scenes. We obtain these trajectories using dataset-specific auxiliary modalities: sensor depth, LiDAR, SLAM and SfM points, human meshes, posed object meshes, and simulation. Every sequence and trajectory is visually verified by human annotators. Across more than 30 baselines, no method comes close to solving the task. Surprisingly, existing multi-view point trackers do not consistently outperform monocular point trackers. By evaluating reconstruction and point tracking on the same datasets, TAPVid-MV helps distinguish errors in recovered geometry from errors in point correspondence. Through this joint analysis, we identify geometry recovery as a major bottleneck for accurate 3D point tracking. Beyond multi-view 3D point tracking, our released annotations support monocular 2D and 3D point tracking, future-trajectory prediction, and 4D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。