arXiv:2510.20951cs.CV2025-10

用生成式方法追踪视频中多模式点轨迹,提升遮挡下的跟踪精度。

Generative Point Tracking with Flow Matching

  • 基于流匹配的生成框架,融合窗口依赖先验与坐标方差调度。
  • 在遮挡点上实现最先进精度,可见点性能也保持竞争力。
  • 适合需要多模态轨迹建模的视频跟踪任务,如复杂场景追踪。

由于外观变化和遮挡等视觉模糊导致的不确定性,视频中点的跟踪极具挑战性。尽管现有最优判别模型能在遮挡情况下回归长期点轨迹,但它们仅能预测单一均值(或众数),无法捕捉多模态特性。为此,我们提出生成式点追踪器(GenPT),一种建模多模态轨迹的生成框架。GenPT采用新颖的流匹配训练方式,结合判别式追踪器的迭代优化、跨窗口一致性先验以及专为点坐标设计的方差调度。我们展示如何在推理阶段利用生成样本,通过基于模型置信度的最优优先搜索策略改进轨迹估计。在PointOdyssey、Dynamic Replica和TAP-Vid标准基准上评估了GenPT,并引入含额外遮挡的TAP-Vid变体以评估遮挡点跟踪性能。结果表明,GenPT能够有效捕捉点轨迹的多模态性,在遮挡点上达到最先进的跟踪精度,同时在可见点上的表现与现有判别式追踪器相当。

原文摘要 · Abstract (English)

Tracking a point through a video can be a challenging task due to uncertainty arising from visual obfuscations, such as appearance changes and occlusions. Although current state-of-the-art discriminative models excel in regressing long-term point trajectory estimates -- even through occlusions -- they are limited to regressing to a mean (or mode) in the presence of uncertainty, and fail to capture multi-modality. To overcome this limitation, we introduce Generative Point Tracker (GenPT), a generative framework for modelling multi-modal trajectories. GenPT is trained with a novel flow matching formulation that combines the iterative refinement of discriminative trackers, a window-dependent prior for cross-window consistency, and a variance schedule tuned specifically for point coordinates. We show how our model's generative capabilities can be leveraged to improve point trajectory estimates by utilizing a best-first search strategy on generated samples during inference, guided by the model's own confidence of its predictions. Empirically, we evaluate GenPT against the current state of the art on the standard PointOdyssey, Dynamic Replica, and TAP-Vid benchmarks. Further, we introduce a TAP-Vid variant with additional occlusions to assess occluded point tracking performance and highlight our model's ability to capture multi-modality. GenPT is capable of capturing the multi-modality in point trajectories, which translates to state-of-the-art tracking accuracy on occluded points, while maintaining competitive tracking accuracy on visible points compared to extant discriminative point trackers.

点追踪生成模型多模态视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。