arXiv:2512.02339cs.CVcs.AI2025-12NeurIPS

利用视频扩散模型的运动表征,无监督追踪相似物体表现更优

Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision

  • 基于扩散模型去噪阶段的运动表征进行无监督跟踪
  • 在相似物体追踪任务上比现有方法提升6个百分点
  • 适合需要少标注数据的视觉追踪场景

区分外观相似物体的运动仍是计算机视觉中的关键挑战。尽管有监督追踪器表现良好,但当前自监督追踪器在视觉线索模糊时表现不佳,限制了其可扩展性和泛化能力。我们发现预训练视频扩散模型天然学习到适用于追踪的运动表征,无需任务特定训练。这一能力源于其去噪过程在早期高噪声阶段分离出运动信息,与后期外观优化阶段相区别。基于此,我们的自监督追踪器显著提升了区分外观相似物体的能力,这是现有方法未充分探索的短板。该方法在主流基准和新引入的相似物体追踪测试中,相比近期自监督方法最高提升6个百分点。可视化结果表明,这些扩散模型衍生的运动表征能有效应对视角变化和形变等挑战,实现对完全相同物体的稳定追踪。

原文摘要 · Abstract (English)

Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labeled data. We find that pre-trained video diffusion models inherently learn motion representations suitable for tracking without task-specific training. This ability arises because their denoising process isolates motion in early, high-noise stages, distinct from later appearance refinement. Capitalizing on this discovery, our self-supervised tracker significantly improves performance in distinguishing visually similar objects, an underexplored failure point for existing methods. Our method achieves up to a 6-point improvement over recent self-supervised approaches on established benchmarks and our newly introduced tests focused on tracking visually similar items. Visualizations confirm that these diffusion-derived motion representations enable robust tracking of even identical objects across challenging viewpoint changes and deformations.

视频生成扩散模型目标追踪自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。