通过渐进式训练提升视觉目标跟踪精度,效果优于现有方法。
Progressive Scaling Visual Object Tracking
- 采用渐进式训练策略,分阶段扩大数据量、模型规模和输入分辨率。
- 在多个基准上超越当前最优方法,显著提升跟踪准确率与泛化能力。
- 适用于多种视觉任务,具有良好的可迁移性,适合追求高精度的跟踪研究者。
本文提出一种面向视觉目标跟踪的渐进式训练策略,系统分析了训练数据量、模型规模和输入分辨率对跟踪性能的影响。实证研究表明,虽然单独扩大任一因素均能显著提升跟踪精度,但直接训练存在优化不足和迭代优化受限的问题。为此,我们引入DT-Training框架,融合小教师迁移与双分支对齐机制,充分释放模型潜力。所得到的扩展跟踪器在多个基准上持续优于当前最优方法,展现出强大的泛化能力和可迁移性。此外,我们验证了该方法在其他任务中的广泛适用性,凸显其超越跟踪任务的通用价值。
原文摘要 · Abstract (English)
In this work, we propose a progressive scaling training strategy for visual object tracking, systematically analyzing the influence of training data volume, model size, and input resolution on tracking performance. Our empirical study reveals that while scaling each factor leads to significant improvements in tracking accuracy, naive training suffers from suboptimal optimization and limited iterative refinement. To address this issue, we introduce DT-Training, a progressive scaling framework that integrates small teacher transfer and dual-branch alignment to maximize model potential. The resulting scaled tracker consistently outperforms state-of-the-art methods across multiple benchmarks, demonstrating strong generalization and transferability of the proposed method. Furthermore, we validate the broader applicability of our approach to additional tasks, underscoring its versatility beyond tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。