单目视频3D点跟踪新方法,速度超快精度高。
SpatialTrackerV2: 3D Point Tracking Made Easy
- 统一3D追踪、深度估计与相机位姿,端到端可训练。
- 在多种数据上联合学习,精度比现有方法高30%。
- 适合需要高速3D追踪的实时应用,如AR/VR。
我们提出SpatialTrackerV2,一种基于单目视频的前馈式3D点追踪方法。不同于依赖现成组件的模块化流水线,该方法将点追踪、单目深度估计和相机位姿估计之间的内在关联统一为一个高性能的前馈3D点追踪器。其将世界空间中的3D运动分解为场景几何、相机自运动和像素级物体运动,并采用完全可微分的端到端架构,支持在多种数据集(包括合成序列、带姿态的RGB-D视频和无标注真实场景视频)上进行可扩展训练。通过联合学习异构数据中的几何与运动信息,SpatialTrackerV2在性能上超越现有3D追踪方法30%,同时达到领先动态3D重建方法的精度,且运行速度提升50倍。
原文摘要 · Abstract (English)
We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50$\times$ faster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。