arXiv:2601.04777cs.CVcs.AI2026-01AAAI

用大模型实现跨图像通用视觉定位,支持多种复杂任务。

GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models

  • 构建统一框架与新数据集,支持多图像间跨线索推理。
  • 在多图像定位任务上提升9.7%,单图像定位提升9.1%。
  • 适合需要跨图理解与复杂推理的视觉语言研究者。

多模态大模型在单图定位和多图理解方面已取得显著进展。近期虽有方法尝试解决多图定位问题,但受限于单一目标定位和任务类型有限,主要因缺乏对通用定位任务的统一建模。为此,我们提出 GeM-VG,一种具备通用多图像视觉定位能力的多模态大模型。我们系统地按跨图线索依赖与推理复杂度对现有任务进行分类,并引入 MG-Data-240K 数据集,弥补了现有数据集在目标数量和图像关系上的不足。为应对多样化任务的鲁棒性挑战,我们设计了一种融合思维链(CoT)推理与直接回答的混合强化微调策略,采用基于规则的奖励机制引导类似 R1 的算法,有效增强模型感知与推理能力。大量实验表明,该模型在多图像定位任务上分别优于先前领先模型 2.0% 和 9.7%(在 MIG-Bench 与 MC-Bench 上),在单图像定位任务上较基线提升 9.1%(在 ODINW 上)。此外,模型仍保持强大的通用多图像理解能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated impressive progress in single-image grounding and general multi-image understanding. Recently, some methods begin to address multi-image grounding. However, they are constrained by single-target localization and limited types of practical tasks, due to the lack of unified modeling for generalized grounding tasks. Therefore, we propose GeM-VG, an MLLM capable of Generalized Multi-image Visual Grounding. To support this, we systematically categorize and organize existing multi-image grounding tasks according to their reliance of cross-image cues and reasoning, and introduce the MG-Data-240K dataset, addressing the limitations of existing datasets regarding target quantity and image relation. To tackle the challenges of robustly handling diverse multi-image grounding tasks, we further propose a hybrid reinforcement finetuning strategy that integrates chain-of-thought (CoT) reasoning and direct answering, considering their complementary strengths. This strategy adopts an R1-like algorithm guided by a carefully designed rule-based reward, effectively enhancing the model's overall perception and reasoning capabilities. Extensive experiments demonstrate the superior generalized grounding capabilities of our model. For multi-image grounding, it outperforms the previous leading MLLMs by 2.0% and 9.7% on MIG-Bench and MC-Bench, respectively. In single-image grounding, it achieves a 9.1% improvement over the base model on ODINW. Furthermore, our model retains strong capabilities in general multi-image understanding.

多图像定位大模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。