四模态融合提升复杂场景追踪鲁棒性
Towards General Multimodal Visual Tracking
- 设计多尺度扫描的QuadFusion模型,高效融合RGB、热成像、事件与语言四模态信息
- 在600个视频序列(384.7万帧)上实现比现有方法更高的追踪精度与稳定性
- 适用于复杂环境下的通用多模态追踪,尤其适合高要求工业或安防场景
现有跨模态追踪研究多集中于双模态场景,如RGB-热成像、RGB-事件、RGB-语言。尽管通过融合互补信息取得良好效果,但在复杂场景下仍受限。本文提出通用多模态视觉追踪任务,全面利用RGB、热红外、事件和语言四种模态,在挑战性条件下实现鲁棒追踪。为此构建了大规模高质量基准QuadTrack600,包含600个视频序列,共384.7万帧高分辨率(640x480)帧组,四模态空间对齐并精细标注边界框,提供21种序列级挑战属性用于深入分析。针对四模态信息量差异大、计算负担重的问题,提出QuadFusion方法,采用多尺度融合Mamba结构,以四种不同扫描尺度实现充分交互,克服指数级计算开销。在QuadTrack600及三个双模态数据集(LasHeR、VisEvent、TNL2K)上的实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Existing multimodal tracking studies focus on bi-modal scenarios such as RGB-Thermal, RGB-Event, and RGB-Language. Although promising tracking performance is achieved through leveraging complementary cues from different sources, it remains challenging in complex scenes due to the limitations of bi-modal scenarios. In this work, we introduce a general multimodal visual tracking task that fully exploits the advantages of four modalities, including RGB, thermal infrared, event, and language, for robust tracking under challenging conditions. To provide a comprehensive evaluation platform for general multimodal visual tracking, we construct QuadTrack600, a large-scale, high-quality benchmark comprising 600 video sequences (totaling 384.7K high-resolution (640x480) frame groups). In each frame group, all four modalities are spatially aligned and meticulously annotated with bounding boxes, while 21 sequence-level challenge attributes are provided for detailed performance analysis. Despite quad-modal data provides richer information, the differences in information quantity among modalities and the computational burden from four modalities are two challenging issues in fusing four modalities. To handle these issues, we propose a novel approach called QuadFusion, which incorporates an efficient Multiscale Fusion Mamba with four different scanning scales to achieve sufficient interactions of the four modalities while overcoming the exponential computational burden, for general multimodal visual tracking. Extensive experiments on the QuadTrack600 dataset and three bi-modal tracking datasets, including LasHeR, VisEvent, and TNL2K, validate the effectiveness of our QuadFusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。