arXiv:2501.03220cs.CV2025-01被引 4

融合全局语义与局部时序特征,实现视频中任意点的精准长期跟踪。

ProTracker: Probabilistic Integration for Robust and Accurate Point Tracking

  • 结合局部光流预测与全局热图观测,构建概率化跟踪框架。
  • 在多个基准上超越现有优化方法,达到当前最优性能。
  • 适合需要高精度、抗遮挡长时跟踪的应用场景。

我们提出 ProTracker,一种用于视频中任意点长时间密集跟踪的新框架。以往依赖全局代价体积的方法虽能有效处理大范围遮挡和场景变化,但精度与时间感知能力不足;而基于局部迭代的方法虽能精确追踪平滑变换的场景,却难以应对遮挡与漂移问题。为此,我们设计了一种概率框架,利用局部光流进行预测,并通过优化的全局热图实现观测,有效融合了全局语义信息与具有时间感知的低层特征,实现了视频中任意点的高精度、强鲁棒性长期跟踪。大量实验证明,ProTracker 在基于优化的方法中达到最先进水平,并在多个基准上超越监督前馈方法。代码与模型将在发表后公开。

原文摘要 · Abstract (English)

We propose ProTracker, a novel framework for accurate and robust long-term dense tracking of arbitrary points in videos. Previous methods relying on global cost volumes effectively handle large occlusions and scene changes but lack precision and temporal awareness. In contrast, local iteration-based methods accurately track smoothly transforming scenes but face challenges with occlusions and drift. To address these issues, we propose a probabilistic framework that marries the strengths of both paradigms by leveraging local optical flow for predictions and refined global heatmaps for observations. This design effectively combines global semantic information with temporally aware low-level features, enabling precise and robust long-term tracking of arbitrary points in videos. Extensive experiments demonstrate that ProTracker attains state-of-the-art performance among optimization-based approaches and surpasses supervised feed-forward methods on multiple benchmarks. The code and model will be released after publication.

点跟踪视频理解概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。