融合可见光与热成像,实现全天候精准目标追踪。
RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking
- 构建多模态融合框架RTrack,整合视觉、热成像与文本信息。
- 在388条描述、1250个目标的数据集上实现高精度追踪。
- 适合关注夜视、烟雾等复杂场景追踪的研究者。
由于具备人机交互友好特性,指代式多目标追踪受到越来越多关注,但在夜间、烟雾等低可见度条件下表现受限。为克服此问题,本文提出全新的RGB-热成像指代式多目标追踪任务(RT-RMOT),旨在融合可见光外观特征与热成像的光照鲁棒性,实现全天候指代追踪。为推动该方向研究,我们构建首个基于RGB-热成像模态的指代多目标追踪数据集RefRT,包含388条语言描述、1250个被追踪目标和166,147个语言-可见光-热成像(L-RGB-T)三元组。同时提出RTrack框架,基于多模态大语言模型(MLLM)融合RGB、热成像与文本特征。针对初始框架优化空间,引入分组序列策略优化(GSPO)以挖掘模型潜力;为缓解强化学习微调中的训练不稳定性,设计截断优势缩放(CAS)策略抑制梯度爆炸;并设计结构化输出奖励与综合检测奖励,平衡探索与利用,提升目标感知的完整性和准确性。在RefRT数据集上的大量实验验证了RTrack框架的有效性。
原文摘要 · Abstract (English)
Referring Multi-Object Tracking has attracted increasing attention due to its human-friendly interactive characteristics, yet it exhibits limitations in low-visibility conditions, such as nighttime, smoke, and other challenging scenarios. To overcome this limitation, we propose a new RGB-Thermal RMOT task, named RT-RMOT, which aims to fuse RGB appearance features with the illumination robustness of the thermal modality to enable all-day referring multi-object tracking. To promote research on RT-RMOT, we construct the first Referring Multi-Object Tracking dataset under RGB-Thermal modality, named RefRT. It contains 388 language descriptions, 1,250 tracked targets, and 166,147 Language-RGB-Thermal (L-RGB-T) triplets. Furthermore, we propose RTrack, a framework built upon a multimodal large language model (MLLM) that integrates RGB, thermal, and textual features. Since the initial framework still leaves room for improvement, we introduce a Group Sequence Policy Optimization (GSPO) strategy to further exploit the model's potential. To alleviate training instability during RL fine-tuning, we introduce a Clipped Advantage Scaling (CAS) strategy to suppress gradient explosion. In addition, we design Structured Output Reward and Comprehensive Detection Reward to balance exploration and exploitation, thereby improving the completeness and accuracy of target perception. Extensive experiments on the RefRT dataset demonstrate the effectiveness of the proposed RTrack framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。