arXiv:2501.03616cs.CV2025-01被引 4

通过双模板与动态消融提升红外+可见光跟踪精度

BTMTrack: Robust RGB-T Tracking via Dual-template Bridging and Temporal-Modal Candidate Elimination

论文配图:BTMTrack: Robust RGB-T Tracking via Dual-template Bridging and Temporal-Modal Candidate Elimination
图 1 · 摘自论文原文
  • 用双模板网络融合时序信息,增强目标表征
  • 在LasHeR数据集上达72.3%准确率,优于现有方法
  • 适合复杂光照与天气下的多模态目标跟踪

RGB-T跟踪利用可见光与热成像的互补优势,应对低照度和恶劣天气等挑战。现有方法常难以有效整合时序信息并高效进行跨模态交互,限制了对动态目标的适应能力。本文提出BTMTrack框架,核心为双模板骨干网络与时间-模态候选消除(TMCE)策略。双模板骨干网络有效融合时序信息,而TMCE通过评估时序与模态相关性,聚焦目标相关特征,降低计算开销并抑制背景噪声。在此基础上,设计时序双模板桥接(TDTB)模块,通过动态过滤的特征实现精准跨模态融合,强化模板与搜索区域的交互。在三个基准数据集上的实验表明,该方法性能领先:在LasHeR测试集上达到72.3%的精度,在RGBT210和RGBT234数据集上也表现优异。

原文摘要 · Abstract (English)

RGB-T tracking leverages the complementary strengths of RGB and thermal infrared (TIR) modalities to address challenging scenarios such as low illumination and adverse weather. However, existing methods often fail to effectively integrate temporal information and perform efficient cross-modal interactions, which constrain their adaptability to dynamic targets. In this paper, we propose BTMTrack, a novel framework for RGB-T tracking. The core of our approach lies in the dual-template backbone network and the Temporal-Modal Candidate Elimination (TMCE) strategy. The dual-template backbone effectively integrates temporal information, while the TMCE strategy focuses the model on target-relevant tokens by evaluating temporal and modal correlations, reducing computational overhead and avoiding irrelevant background noise. Building upon this foundation, we propose the Temporal Dual Template Bridging (TDTB) module, which facilitates precise cross-modal fusion through dynamically filtered tokens. This approach further strengthens the interaction between templates and the search region. Extensive experiments conducted on three benchmark datasets demonstrate the effectiveness of BTMTrack. Our method achieves state-of-the-art performance, with a 72.3% precision rate on the LasHeR test set and competitive results on RGBT210 and RGBT234 datasets.

多模态跟踪红外视觉时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。