让大模型学会理解红外、深度等多模态图像,提升复杂场景感知能力。
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
- 用视觉思维链引导模型跨模态推理,从RGB拓展到红外等多模态
- 两阶段训练使模型在三类任务上比基线高22.71%的准确率
- 首个针对多模态定位的基准数据集,适合多模态视觉研究者
多模态大语言模型(MLLM)主要基于RGB模态预训练,限制了其在红外、深度、事件等关键模态上的表现。为此,我们提出RGBX-R1框架,增强MLLM在各类X视觉模态下的感知与推理能力。采用理解-关联-验证(UAV)提示策略构建视觉模态思维链(VM-CoT),将模型对RGB的理解扩展至多模态。设计两阶段训练:冷启动监督微调(CS-SFT)利用VM-CoT指导推理过程,建立基础模态认知;在此基础上,基于GRPO改进的时空强化微调(ST-RFT)引入模态理解时空奖励(MuST),强化模态推理。我们构建了首个RGBX定位基准,实验表明在三个任务上相较基线提升22.71%。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLM) are primarily pre-trained on the RGB modality, thereby limiting their performance on other modalities, such as infrared, depth, and event data, which are crucial for complex scenarios. To address this, we propose RGBX-R1, a framework to enhance MLLM's perception and reasoning capacities across various X visual modalities. Specifically, we employ an Understand-Associate-Validate (UAV) prompting strategy to construct the Visual Modality Chain-of-Thought (VM-CoT), which aims to expand the MLLMs' RGB understanding capability into X modalities. To progressively enhance reasoning capabilities, we introduce a two-stage training paradigm: Cold-Start Supervised Fine-Tuning (CS-SFT) and Spatio-Temporal Reinforcement Fine-Tuning (ST-RFT). CS-SFT supervises the reasoning process with the guidance of VM-CoT, equipping the MLLM with fundamental modality cognition. Building upon GRPO, ST-RFT employs a Modality-understanding Spatio-Temporal (MuST) reward to reinforce modality reasoning. Notably, we construct the first RGBX-Grounding benchmark, and extensive experiments verify our superiority in multimodal understanding and spatial perception, outperforming baselines by 22.71% on three RGBX grounding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。