用强化学习让多模态模型更懂图像,提升视觉理解能力。
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
- 通过注意力机制识别视觉依赖句段并掩码,引导模型基于图像推理重建。
- 在多个零样本任务中表现提升,微调后鲁棒性显著增强。
- 首个将强化学习直接用于大模型预训练的方法,适合视觉理解研究者。
多模态预训练受限于图像-标题对的描述偏差,导致模型偏好表面语言线索而非基于图像的理解。我们提出MMRPT,一种基于掩码视觉依赖推理的多模态强化预训练框架,强化大型多模态语言模型的视觉推理能力。首次将强化学习直接引入大视觉-语言模型的预训练过程,使学习信号奖励视觉锚定而非标题模仿。MMRPT通过视觉标记注意力估计句子级视觉依赖性,并掩码高视觉依赖段落;模型在语义-视觉奖励引导下,通过视觉锚定推理重建这些片段。实验表明,该方法在多种基准上实现一致的零样本提升,并在监督微调下显著增强鲁棒性,证明强化驱动的掩码推理为多模态模型提供了更可靠、泛化性更强的预训练目标。
原文摘要 · Abstract (English)
Multimodal pre-training remains constrained by the descriptive bias of image-caption pairs, leading models to favor surface linguistic cues over grounded visual understanding. We introduce MMRPT, a masked multimodal reinforcement pre-training framework that strengthens visual reasoning in MLLMs. We are the first to incorporate reinforcement learning directly into the pre-training of large vision-language models, enabling learning signals that reward visual grounding rather than caption imitation. MMRPT constructs masked multimodal data by estimating sentence-level visual dependency via attention over visual tokens and masking highly vision-dependent segments; the model reconstructs these spans through vision-grounded reasoning guided by a semantic-visual reward. Experiments show consistent zero-shot gains across diverse benchmarks and substantially improved robustness under supervised fine-tuning, demonstrating that reinforcement-driven masked reasoning provides a more reliable and generalizable pre-training objective for multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。