arXiv:2410.11831cs.CV2024-10被引 98

用真实视频伪标签训练,让点追踪更准更轻量

CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

  • 用现成教师模型生成真实视频伪标签,实现半监督训练
  • 仅需千分之一数据量就超越以往方法,模型更小更简单
  • 适合需要低资源、高鲁棒性点追踪的视觉任务

当前最先进的点追踪模型多在合成数据上训练,因真实视频标注困难。这导致合成与真实视频间存在统计差异,影响性能。为此,本文提出CoTracker3,包含新追踪模型和半监督训练方案,利用现成教师模型为无标注真实视频生成伪标签,实现高效训练。新模型简化或移除了前代追踪器中的复杂组件,架构更轻量。该训练方式远比之前简单,且仅用1,000倍少的数据即取得更好效果。研究还分析了增加真实无监督数据对模型性能的影响。模型提供在线与离线两种版本,能可靠追踪可见及被遮挡点。

原文摘要 · Abstract (English)

Most state-of-the-art point trackers are trained on synthetic data due to the difficulty of annotating real videos for this task. However, this can result in suboptimal performance due to the statistical gap between synthetic and real videos. In order to understand these issues better, we introduce CoTracker3, comprising a new tracking model and a new semi-supervised training recipe. This allows real videos without annotations to be used during training by generating pseudo-labels using off-the-shelf teachers. The new model eliminates or simplifies components from previous trackers, resulting in a simpler and often smaller architecture. This training scheme is much simpler than prior work and achieves better results using 1,000 times less data. We further study the scaling behaviour to understand the impact of using more real unsupervised data in point tracking. The model is available in online and offline variants and reliably tracks visible and occluded points.

点追踪半监督伪标签轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。