提出MTNet,用Transformer实现多模态感知的红外可见光跟踪。
MTNet: Learning modality-aware representation with transformer for RGBT tracking
- 设计模态感知网络,分离并强化可见光与红外特征
- 在三个基准上超越现有方法,且达到实时速度
- 适合需要高精度跨模态目标追踪的研究者
学习鲁棒的多模态表示在RGBT跟踪发展中起关键作用。然而,传统的融合范式和不变的跟踪模板限制了特征交互。本文提出基于Transformer的模态感知跟踪器MTNet。具体地,提出模态感知网络以挖掘模态特异性线索,包含通道聚合与分布模块(CADM)和空间相似性感知模块(SSPM)。随后引入Transformer融合网络捕捉全局依赖,增强实例表示。为精确估计位置并应对尺度变化与形变等挑战,设计三叉预测头与动态更新策略,共同维护可靠模板以促进帧间通信。大量实验表明,所提方法在三个RGBT基准上优于当前最先进模型,同时达到实时速度。
原文摘要 · Abstract (English)
The ability to learn robust multi-modality representation has played a critical role in the development of RGBT tracking. However, the regular fusion paradigm and the invariable tracking template remain restrictive to the feature interaction. In this paper, we propose a modality-aware tracker based on transformer, termed MTNet. Specifically, a modality-aware network is presented to explore modality-specific cues, which contains both channel aggregation and distribution module(CADM) and spatial similarity perception module (SSPM). A transformer fusion network is then applied to capture global dependencies to reinforce instance representations. To estimate the precise location and tackle the challenges, such as scale variation and deformation, we design a trident prediction head and a dynamic update strategy which jointly maintain a reliable template for facilitating inter-frame communication. Extensive experiments validate that the proposed method achieves satisfactory results compared with the state-of-the-art competitors on three RGBT benchmarks while reaching real-time speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。