全面梳理通用目标追踪方法,重点解析基于Transformer的新进展。
A Deep Dive into Generic Object Tracking: A Survey
- 按思路将追踪方法分为孪生、判别与Transformer三类,系统对比设计原理。
- 指出基于Transformer的方法因时空建模强而发展迅猛,性能显著提升。
- 适合想快速掌握追踪领域全貌或关注Transformer应用的研究者。
通用目标追踪在计算机视觉中仍具重要性但面临挑战,尤其在遮挡、相似干扰物和外观变化等复杂时空动态下。过去二十年间,涌现了包括孪生网络、判别式追踪器以及近年兴起的Transformer-based方法在内的多种范式。现有综述或聚焦单一类别,或泛泛覆盖多类,本文则对三类方法进行全方位回顾,特别强调快速发展的Transformer方法。通过定性与定量比较,分析各类方法的核心设计、创新点与局限性。研究提出新颖分类体系,提供统一的可视化与表格对比,并从多角度组织现有追踪器,总结主流评估基准,凸显基于Transformer的追踪方法因具备强大时空建模能力而取得的快速进展。
原文摘要 · Abstract (English)
Generic object tracking remains an important yet challenging task in computer vision due to complex spatio-temporal dynamics, especially in the presence of occlusions, similar distractors, and appearance variations. Over the past two decades, a wide range of tracking paradigms, including Siamese-based trackers, discriminative trackers, and, more recently, prominent transformer-based approaches, have been introduced to address these challenges. While a few existing survey papers in this field have either concentrated on a single category or widely covered multiple ones to capture progress, our paper presents a comprehensive review of all three categories, with particular emphasis on the rapidly evolving transformer-based methods. We analyze the core design principles, innovations, and limitations of each approach through both qualitative and quantitative comparisons. Our study introduces a novel categorization and offers a unified visual and tabular comparison of representative methods. Additionally, we organize existing trackers from multiple perspectives and summarize the major evaluation benchmarks, highlighting the fast-paced advancements in transformer-based tracking driven by their robust spatio-temporal modeling capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。