arXiv:2511.21946cs.CV2025-11NeurIPS

让视频模型追踪全景中超出视野的点,突破传统视觉局限。

TAPVid-360: Tracking Any Point in 360 from Narrow Field of View Video

  • 用360度视频生成窄视角数据,通过2D跟踪计算真值方向
  • 在10000个视频上实现跨视域点追踪,性能超越现有方法
  • 适合研究全景感知、动态场景建模的开发者与研究人员

人类能构建周围环境的全景认知模型,保持物体恒常性并推断可见区域外的场景结构。而当前人工视觉系统在持续性全景理解方面表现不足,通常仅以自我中心视角逐帧处理场景。这一局限在「追踪任意点」(TAP)任务中尤为明显,现有方法无法追踪视野外的2D点。为此,我们提出TAPVid-360新任务:要求模型预测视频序列中查询场景点的3D方向,即使该点远在窄视角之外。该任务促进学习非中心化的场景表征,无需依赖动态4D真实场景模型进行训练。我们利用360视频作为监督信号,将其重采样为窄视角视角,并通过2D跟踪管道计算点的真值方向。我们构建了新数据集与基准TAPVid360-10k,包含10,000个带有真值方向追踪的视角视频。基线模型将CoTracker v3适配为逐点预测旋转以更新方向,优于现有TAP和TAPVid 3D方法。

原文摘要 · Abstract (English)

Humans excel at constructing panoramic mental models of their surroundings, maintaining object permanence and inferring scene structure beyond visible regions. In contrast, current artificial vision systems struggle with persistent, panoramic understanding, often processing scenes egocentrically on a frame-by-frame basis. This limitation is pronounced in the Track Any Point (TAP) task, where existing methods fail to track 2D points outside the field of view. To address this, we introduce TAPVid-360, a novel task that requires predicting the 3D direction to queried scene points across a video sequence, even when far outside the narrow field of view of the observed video. This task fosters learning allocentric scene representations without needing dynamic 4D ground truth scene models for training. Instead, we exploit 360 videos as a source of supervision, resampling them into narrow field-of-view perspectives while computing ground truth directions by tracking points across the full panorama using a 2D pipeline. We introduce a new dataset and benchmark, TAPVid360-10k comprising 10k perspective videos with ground truth directional point tracking. Our baseline adapts CoTracker v3 to predict per-point rotations for direction updates, outperforming existing TAP and TAPVid 3D methods. Project page: https://finlay-hudson.github.io/tapvid360

全景视觉点追踪360视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。