arXiv:2507.00748cs.CV2025-07被引 10

用强化学习提升多图理解能力,让大模型更会推理。

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

  • 先生成思维链数据,再用低秩微调优化模型
  • 在多图基准上提升9.04%,跨领域平均增4.41%
  • 适合需要复杂视觉推理的AI研究者

多模态大语言模型在单图视觉定位任务中表现良好,但在需要跨图推理和多模态指令的现实任务中表现不佳。为此,本文提出一种基于强化学习(RL)的后训练策略,用于改善多图接地任务中的表现。首先合成高质量的思维链(CoT)数据以实现冷启动初始化,随后使用低秩适配(LoRA)进行监督微调(SFT)。接着,通过合并后的SFT模型进行拒绝采样,构建可靠的强化学习数据,并采用基于规则的强化学习引导模型走向最优推理路径。大量实验表明,该方法在MIG-Bench上提升9.04%,在七个域外基准上平均提升4.41%。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instructions. To address this, we adopt a reinforcement learning (RL) based post-training strategy for MLLMs in multi-image grounding tasks. We first synthesize high-quality chain-of-thought (CoT) data for cold-start initialization, followed by supervised fine-tuning (SFT) using low-rank adaptation (LoRA). Subsequently, we apply rejection sampling with the merged SFT model to curate reliable RL data and use rule-based RL to guide the model toward optimal reasoning paths. Extensive experiments demonstrate the effectiveness of our approach, achieving +9.04% on MIG-Bench and +4.41% on average across seven out-of-domain benchmarks.

多图理解强化学习视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。