arXiv:2603.02573cs.CV2026-03被引 7

实现单目视频中所有像素的高效世界坐标系3D轨迹追踪

Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels

  • 基于ViT构建全局3D场景表征,用新型3D相关性同步估计任意帧间像素2D/3D光流
  • 在多个基准上超越现有方法,在2D/3D光流与3D追踪任务中均表现最优
  • 适合需要高精度4D重建的现实场景应用,如自动驾驶与虚拟现实

从单目视频中估计每个像素的3D轨迹对于全面理解视频的3D动态至关重要。现有单目3D追踪方法要么仅能追踪首帧稀疏点,要么采用耗时的优化框架进行稠密追踪。本文提出一种前馈模型Track4World,实现世界坐标系下所有像素的高效整体3D追踪。该模型基于VGGT风格的ViT构建全局3D场景表征,引入新颖的3D相关性机制,可同时估计任意帧对间的像素级2D与3D稠密光流。结合重建的3D几何信息,实现后续高效的全像素3D追踪。大量实验表明,本方法在多个基准上持续优于现有方法,在2D/3D光流估计与3D追踪任务中均表现优异,展现出对真实世界4D重建任务的鲁棒性与可扩展性。

原文摘要 · Abstract (English)

Estimating the 3D trajectory of every pixel from a monocular video is crucial and promising for a comprehensive understanding of the 3D dynamics of videos. Recent monocular 3D tracking works demonstrate impressive performance, but are limited to either tracking sparse points on the first frame or a slow optimization-based framework for dense tracking. In this paper, we propose a feedforward model, called Track4World, enabling an efficient holistic 3D tracking of every pixel in the world-centric coordinate system. Built on the global 3D scene representation encoded by a VGGT-style ViT, Track4World applies a novel 3D correlation scheme to simultaneously estimate the pixel-wise 2D and 3D dense flow between arbitrary frame pairs. The estimated scene flow, along with the reconstructed 3D geometry, enables subsequent efficient 3D tracking of every pixel of this video. Extensive experiments on multiple benchmarks demonstrate that our approach consistently outperforms existing methods in 2D/3D flow estimation and 3D tracking, highlighting its robustness and scalability for real-world 4D reconstruction tasks.

3D追踪单目视频场景流视觉4D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。