融合可见光、深度与热成像三模态信号,提升复杂场景追踪性能
Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and Algorithm
- 通过正交投影约束融合深度与热成像,作为提示输入预训练视觉模型
- 在RGBDT500数据集上,追踪准确率显著优于双模态方法
- 适合需要高鲁棒性追踪的自动驾驶与安防应用
现有多模态目标追踪方法主要聚焦于双模态(如RGB-深度或RGB-热红外),在复杂场景下仍受限于输入模态数量。为此,本文提出一种新多模态追踪任务,融合可见光RGB、深度(D)和热红外(TIR)三种互补模态,以增强复杂场景下的鲁棒性。为此构建了新的多模态追踪数据集RGBDT500,包含500段视频,三模态帧同步且空间对齐,每帧均有精确的目标边界框标注。同时提出新型多模态追踪器RDTTrack:基于预训练的仅含RGB的追踪模型,利用提示学习技术融合深度与热红外信息。具体而言,通过提出的正交投影约束融合两种模态,再将其作为提示注入预训练基础追踪模型,有效协调三模态互补特征。实验结果表明,该方法在追踪精度与复杂场景鲁棒性方面显著优于现有双模态方法。数据集与源码已公开:https://xuefeng-zhu5.github.io/RGBDT500。
原文摘要 · Abstract (English)
Existing multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three complementary modalities, including visible RGB, Depth (D), and Thermal Infrared (TIR), aiming to enhance robustness in complex scenarios. To support this task, we construct a new multi-modal tracking dataset, coined RGBDT500, which consists of 500 videos with synchronised frames across the three modalities. Each frame provides spatially aligned RGB, depth, and thermal infrared images with precise object bounding box annotations. Furthermore, we propose a novel multi-modal tracker, dubbed RDTTrack. RDTTrack integrates tri-modal information for robust tracking by leveraging a pretrained RGB-only tracking model and prompt learning techniques. In specific, RDTTrack fuses thermal infrared and depth modalities under a proposed orthogonal projection constraint, then integrates them with RGB signals as prompts for the pre-trained foundation tracking model, effectively harmonising tri-modal complementary cues. The experimental results demonstrate the effectiveness and advantages of the proposed method, showing significant improvements over existing dual-modal approaches in terms of tracking accuracy and robustness in complex scenarios. The dataset and source code are publicly available at https://xuefeng-zhu5.github.io/RGBDT500.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。