提出空间-模态联合建模方法,提升复杂光照下多模态目标跟踪性能。
SMAC: Spatial-Modal Joint Modeling and Adaptive Representation Collapse for Multimodal Object Tracking

- 通过解耦3D卷积与幅相分解,联合建模空间与模态特征
- 在RNT模态上达63.31 HOTA和79.21 MOTA,优于多个先进方法
- 适配动态融合策略,适合光照复杂场景下的多模态跟踪应用
复杂光照下多模态多目标跟踪(MOT)仍面临空间与模态特征联合建模不足及固定融合策略适应性差的问题。本文提出一种空间-模态卷积融合与基于知识蒸馏提示的多模态MOT框架。构建空间-模态融合主干网络:Basic模块通过解耦3D卷积实现空间特征提取与模态交互,Mixed模块利用幅相分解建模非线性跨模态相关性。此外,设计表示坍缩网络实现自适应多模态融合:通过教师监督生成动态模态权重的蒸馏提示模块(DPG),以及在表示坍缩过程中保留判别信息的全局模态差异聚合模块(GMDA)。在UniRTL数据集上的大量实验表明,所提方法有效。该追踪器在RNT模态上达到63.31 HOTA和79.21 MOTA,优于多个最先进方法,同时保持良好推理效率。源代码与预训练模型已公开于https://github.com/QitaiSun/SMAC。
原文摘要 · Abstract (English)
Multimodal multi-object tracking (MOT) under complex illumination remains challenging due to insufficient joint modeling of spatial and modal features and the limited adaptability of fixed fusion strategies. To address these issues, this paper proposes a spatial-modal convolution fusion and distillation-prompt-based multimodal MOT framework. A spatial-modal fusion backbone is first constructed, where a Basic module performs spatial feature extraction and modal interaction via decoupled 3D convolution, while a Mixed module models nonlinear cross-modal correlations through amplitude-phase decomposition. In addition, a representation collapse network is designed for adaptive multimodal fusion. A Distillation Prompt Guidance (DPG) module generates dynamic modal weights under teacher supervision, and a Global Modal Difference Aggregation (GMDA) module preserves discriminative information during multimodal representation collapse. Extensive experiments on the UniRTL dataset demonstrate the effectiveness of the proposed method. The proposed tracker achieves 63.31 HOTA and 79.21 MOTA on the RNT modality, outperforming several state-of-the-art methods while maintaining favorable inference efficiency. The source code and pretrained models are publicly available at https://github.com/QitaiSun/SMAC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。