arXiv:2411.15459cs.CV2024-11被引 23

用Mamba模型提升视觉语言追踪的时序建模与特征更新能力

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking

  • 引入时序演化混合状态空间块,实现长序列高效建模
  • 在多个基准上超越现有最优追踪器,准确率提升显著
  • 适合需要多模态动态融合的实时追踪场景

视觉语言追踪任务旨在基于多种模态参考进行目标追踪。现有基于Transformer的方法虽借助自注意力机制的全局建模能力取得显著进展,但在有效利用时序信息和动态更新参考特征方面仍面临挑战。最近,状态空间模型(SSM)中的Mamba展现出高效的长序列建模能力,其状态空间演化过程以线性复杂度具备记忆多模态时序信息的潜力。受此启发,本文提出一种基于Mamba的视觉语言追踪模型——MambaVLT,通过引入时序演化混合状态空间块与选择性局部增强块,捕捉多模态上下文信息并实现自适应参考特征更新。此外,设计了模态选择模块,动态调整视觉与语言参考的权重,缓解单一模态带来的歧义。大量实验表明,该方法在多个基准上均优于当前最优追踪器。

原文摘要 · Abstract (English)

The vision-language tracking task aims to perform object tracking based on various modality references. Existing Transformer-based vision-language tracking methods have made remarkable progress by leveraging the global modeling ability of self-attention. However, current approaches still face challenges in effectively exploiting the temporal information and dynamically updating reference features during tracking. Recently, the State Space Model (SSM), known as Mamba, has shown astonishing ability in efficient long-sequence modeling. Particularly, its state space evolving process demonstrates promising capabilities in memorizing multimodal temporal information with linear complexity. Witnessing its success, we propose a Mamba-based vision-language tracking model to exploit its state space evolving ability in temporal space for robust multimodal tracking, dubbed MambaVLT. In particular, our approach mainly integrates a time-evolving hybrid state space block and a selective locality enhancement block, to capture contextual information for multimodal modeling and adaptive reference feature update. Besides, we introduce a modality-selection module that dynamically adjusts the weighting between visual and language references, mitigating potential ambiguities from either reference type. Extensive experimental results show that our method performs favorably against state-of-the-art trackers across diverse benchmarks.

视觉语言追踪状态空间模型多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。