arXiv:2512.20606cs.CV2025-12被引 3

用视频扩散模型特征提升点追踪的鲁棒性,效果超越传统方法。

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking

  • 采用视频扩散变换器提取时序一致且区分度高的特征。
  • 在合成数据上训练更少迭代次数,仍优于依赖真实数据的CoTracker3。
  • 适合追求高鲁棒性和低实数据依赖的点追踪场景。

尽管现有点追踪方法在标准基准上表现良好,但其特征主干网络通常未针对实际应用所需的时序一致性进行设计。虽然近期工作将视觉基础模型(VFM)特征引入追踪流程,但尚无研究系统分析哪种VFM对点追踪最具鲁棒性。我们首次在零样本设置下,对多种VFM在标准与鲁棒性基准上进行评估。结果表明,视频扩散变换器(DiTs)始终生成最时序连贯且具有区分性的特征,甚至优于在追踪数据上显式训练的ResNet主干。我们推测该优势源于大规模视频预训练、全3D时空注意力机制及扩散训练目标。基于此发现,我们提出DiTracker,通过查询-键匹配代价计算、轻量级ResNet分支的成本融合以及LoRA适配,将视频DiT特征集成到现有追踪框架中。在相同追踪头下,DiTracker仅用合成数据训练更少迭代次数,却超越使用额外真实视频训练的CoTracker3,尤其在挑战性与退化场景下提升显著。它还能跨追踪头泛化并随主干规模扩展,证实生成式视频预训练能提供真实世界先验,降低对大规模真实数据监督的依赖。

原文摘要 · Abstract (English)

Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for robust real-world performance. While recent works incorporate powerful visual foundation model (VFM) features into tracking pipelines, no prior work has systematically analyzed which VFM provides the most robust representations for point tracking. We present the first such analysis, evaluating diverse VFMs in a zero-shot setting on both standard and robustness benchmarks for point tracking. Our study reveals that video diffusion transformers (DiTs) consistently yield the most temporally coherent and discriminative features, even surpassing ResNet backbones explicitly supervised on tracking data. We hypothesize this advantage stem from large-scale video pretraining, full 3D spatio-temporal attention, and a diffusion training objective. Motivated by this finding, we propose DiTracker, which integrates video DiT features into existing tracking frameworks through query-key matching cost computation, cost-level fusion with a lightweight ResNet branch, and LoRA adaptation. Under the same tracking head, DiTracker is trained solely on synthetic data with far fewer iterations, yet outperforms CoTracker3 trained with additional real-world videos, with the largest gains under challenging and corrupted scenarios. It further generalizes across tracking heads and scales with backbone size, confirming that generative video pretraining provides real-world priors that reduce the dependence on large-scale real-data supervision.

点追踪扩散模型视频理解鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。