arXiv:2502.06583cs.CV2025-02被引 23

提出自适应多模态跟踪框架,实现跨模态动态融合

Adaptive Perception for Unified Visual Multi-modal Object Tracking

  • 采用平等建模策略,统一处理多模态输入
  • 在5个数据集上超越现有统一追踪器与专用追踪器
  • 适合需要灵活应对复杂场景的多模态跟踪应用

当前多数多模态跟踪器以RGB为主导,将其他模态视为辅助,并对不同任务分别微调,导致模态依赖失衡,难以在复杂场景中动态利用各模态互补信息,使统一参数模型在多任务中表现不佳。为此,我们提出APTrack,一种面向多模态自适应感知的统一追踪框架。不同于以往方法,APTrack通过平等建模策略构建统一表征,使模型能动态适配不同模态与任务,无需跨任务额外微调。同时,引入可学习令牌的自适应模态交互(AMI)模块,高效实现跨模态信息融合。在五个多模态数据集(RGBT234、LasHeR、VisEvent、DepthTrack、VOT-RGBD2022)上的实验表明,APTrack不仅超越现有先进统一追踪器,还优于针对特定模态设计的追踪器。

原文摘要 · Abstract (English)

Recently, many multi-modal trackers prioritize RGB as the dominant modality, treating other modalities as auxiliary, and fine-tuning separately various multi-modal tasks. This imbalance in modality dependence limits the ability of methods to dynamically utilize complementary information from each modality in complex scenarios, making it challenging to fully perceive the advantages of multi-modal. As a result, a unified parameter model often underperforms in various multi-modal tracking tasks. To address this issue, we propose APTrack, a novel unified tracker designed for multi-modal adaptive perception. Unlike previous methods, APTrack explores a unified representation through an equal modeling strategy. This strategy allows the model to dynamically adapt to various modalities and tasks without requiring additional fine-tuning between different tasks. Moreover, our tracker integrates an adaptive modality interaction (AMI) module that efficiently bridges cross-modality interactions by generating learnable tokens. Experiments conducted on five diverse multi-modal datasets (RGBT234, LasHeR, VisEvent, DepthTrack, and VOT-RGBD2022) demonstrate that APTrack not only surpasses existing state-of-the-art unified multi-modal trackers but also outperforms trackers designed for specific multi-modal tasks.

多模态跟踪自适应感知统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。