用语义提示和噪声学习提升无监督跟踪的上下文建模能力
Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning

- 分阶段引入语义提示与上下文噪声,渐进式增强模型表征能力
- 在无标签视频上实现优于现有方法的跟踪性能,显著提升鲁棒性
- 适合追求高效无监督跟踪的科研与工业应用
从无标签视频中学习鲁棒的上下文知识对推动自监督跟踪至关重要。然而,传统自监督跟踪器缺乏有效的上下文建模能力,而基于非语义查询的上下文关联方法难以适应无标签追踪场景,导致难以学习可靠上下文线索。本文提出一种新型自监督跟踪框架 extbf{ racker},引入双模态上下文关联机制,联合利用细粒度语义提示和上下文噪声,驱动模型学习鲁棒的跟踪表征。遵循由易到难的学习原则,该机制分为两个阶段:早期训练时,向前后向跟踪分支分配实例块标记(提示)以促进跟踪知识获取;随着训练推进,逐步注入上下文噪声干扰特征,促使跟踪器在更复杂的特征空间中学习鲁棒表征。该机制仅在训练阶段使用,保障推理效率。大量实验表明,本方法在无标签视频上显著优于现有方法。
原文摘要 · Abstract (English)
Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective context modeling, while existing context association methods based on non-semantic queries struggle to adapt to unlabeled tracking scenarios, making it difficult to learn reliable contextual cues. In this work, we propose a novel self-supervised tracking framework, named \textbf{\tracker}, which introduces a dual-modal context association mechanism that jointly leverages fine-grained semantic prompts and contextual noise to drive the model toward learning robust tracking representations. Adherent to the easy-to-hard learning principle, our contextual association mechanism operates based on two stages. During early training, instance patch tokens (prompts) are assigned to both forward and backward tracking branches to facilitate the acquisition of tracking knowledge. As training progresses, contextual noise is gradually injected into the model to perturb feature, encouraging the tracker to learn robust tracking representations in a more complex feature space. Thus, this novel contextual association mechanism enables our self-supervised model to learn high-quality tracking representations from unlabeled videos, while being applied exclusively during training to preserve efficient inference. Extensive experiments demonstrate the superiority of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。