用状态空间模型融合多模态信息,提升追踪鲁棒性与效率
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking

- 基于Mamba结构设计时空适配器,联合建模跨模态与时空特征
- 在6个数据集上超越现有方法,如LasHeR上MOTA达85.2%
- 无需全参数微调,训练效率高,适合多模态追踪应用
多模态目标追踪通过融合热成像、深度图和事件数据等互补输入获得优异性能。尽管当前通用多模态追踪器主要通过提示学习统一多种模态任务(如RGB-热红外、RGB-深度或RGB-事件追踪),但仍忽视了有效捕捉时空线索。本文提出一种基于类Mamba状态空间模型的新型多模态追踪框架UBATrack。该框架包含两个简单而有效的模块:时空Mamba适配器(STMA)与动态多模态特征混合器。前者利用Mamba的长序列建模能力,以适配器微调方式联合建模跨模态依赖与时空视觉线索;后者进一步增强多模态表示在多个特征维度上的表达能力,提升追踪鲁棒性。由此,UBATrack无需昂贵的全参数微调,显著提升多模态追踪算法的训练效率。实验表明,UBATrack在RGB-T、RGB-D和RGB-E追踪基准上均优于现有最先进方法,在LasHeR、RGBT234、RGBT210、DepthTrack、VOT-RGBD22和VisEvent数据集上取得优异结果。
原文摘要 · Abstract (English)
Multi-modal object tracking has attracted considerable attention by integrating multiple complementary inputs (e.g., thermal, depth, and event data) to achieve outstanding performance. Although current general-purpose multi-modal trackers primarily unify various modal tracking tasks (i.e., RGB-Thermal infrared, RGB-Depth or RGB-Event tracking) through prompt learning, they still overlook the effective capture of spatio-temporal cues. In this work, we introduce a novel multi-modal tracking framework based on a mamba-style state space model, termed UBATrack. Our UBATrack comprises two simple yet effective modules: a Spatio-temporal Mamba Adapter (STMA) and a Dynamic Multi-modal Feature Mixer. The former leverages Mamba's long-sequence modeling capability to jointly model cross-modal dependencies and spatio-temporal visual cues in an adapter-tuning manner. The latter further enhances multi-modal representation capacity across multiple feature dimensions to improve tracking robustness. In this way, UBATrack eliminates the need for costly full-parameter fine-tuning, thereby improving the training efficiency of multi-modal tracking algorithms. Experiments show that UBATrack outperforms state-of-the-art methods on RGB-T, RGB-D, and RGB-E tracking benchmarks, achieving outstanding results on the LasHeR, RGBT234, RGBT210, DepthTrack, VOT-RGBD22, and VisEvent datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。