arXiv:2603.09287cs.CV2026-03AAAI

针对多模态追踪中融合方式单一、时序信息混淆问题,提出自适应分模态融合与解耦时序传播框架。

Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking

  • 采用专家混合机制,为红外、事件、深度、可见光等模态分配专用处理单元
  • 在五个基准上取得当前最优性能,其中MDTrack S和MDTrack U均领先现有方法
  • 适合需要高精度多模态目标追踪的自动驾驶与机器人应用

现有大多数多模态追踪器采用统一融合策略,忽视了不同模态间的本质差异。同时,它们通过混合令牌传播时序信息,导致时序表征纠缠且缺乏区分性。为此,本文提出MDTrack框架,实现模态感知融合与解耦时序传播。具体而言,针对模态感知融合,为红外、事件、深度和RGB模态分别配置专用专家,利用门控机制动态选择最优专家,实现自适应、模态特异的融合。针对解耦时序传播,引入两个独立的状态空间模型(SSM)结构,分别存储并更新RGB与X模态流的隐藏状态,有效捕捉各自独特的时序特性。为促进两者间协同,设计跨注意力模块在两个SSM输入特征间进行隐式信息交换。最终,时序增强特征通过另一组跨注意力模块融入主干网络,提升对时序信息的利用能力。大量实验表明,所提方法效果显著:在五个多模态追踪基准上,MDTrack S与MDTrack U均达到当前最优性能。

原文摘要 · Abstract (English)

Most existing multimodal trackers adopt uniform fusion strategies, overlooking the inherent differences between modalities. Moreover, they propagate temporal information through mixed tokens, leading to entangled and less discriminative temporal representations. To address these limitations, we propose MDTrack, a novel framework for modality aware fusion and decoupled temporal propagation in multimodal object tracking. Specifically, for modality aware fusion, we allocate dedicated experts to each modality, including infrared, event, depth, and RGB, to process their respective representations. The gating mechanism within the Mixture of Experts dynamically selects the optimal experts based on the input features, enabling adaptive and modality specific fusion. For decoupled temporal propagation, we introduce two separate State Space Model structures to independently store and update the hidden states of the RGB and X modal streams, effectively capturing their distinct temporal information. To ensure synergy between the two temporal representations, we incorporate a set of cross attention modules between the input features of the two SSMs, facilitating implicit information exchange. The resulting temporally enriched features are then integrated into the backbone through another set of cross attention modules, enhancing MDTrack's ability to leverage temporal information. Extensive experiments demonstrate the effectiveness of our proposed method. Both MDTrack S and MDTrack U achieve state of the art performance across five multimodal tracking benchmarks.

多模态追踪融合策略时序建模状态空间模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。