arXiv:2501.05767cs.CLcs.AI2025-01ACL被引 24

首个实现多图自由精准定位的模型,解决复杂场景下视觉定位难题。

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

  • 提出基于思维链的多图定位框架,支持跨图像语义关联。
  • 在多图定位任务中超越现有最佳模型24.94%,媲美70B大模型性能。
  • 开源数据集与评测基准,适合多模态定位研究者使用。

多模态大语言模型(MLLMs)虽已显著提升对单图的细粒度感知和多图整体理解能力,但在复杂多图场景下的精准定位仍面临挑战。为此,我们首先探索将单图定位与多图理解结合的思维链(CoT)框架,虽部分有效,但因非端到端设计而稳定性差,难以捕捉抽象视觉信息。因此,我们提出Migician,首个支持跨多图自由形式、高精度定位的模型。为支持该任务,我们构建了包含63万条样本的MGrounding-630k数据集,涵盖多种多图定位任务及新生成的自由形式指令跟随数据。同时,提出MIG-Bench评测基准,专门评估多图定位能力。实验表明,该模型在多图定位上显著优于现有最佳MLLMs,性能领先24.94%,甚至超越更大规模的70B模型。代码、模型、数据集与评测基准均已在https://migician-vg.github.io/全面开源。

原文摘要 · Abstract (English)

The recent advancement of Multimodal Large Language Models (MLLMs) has significantly improved their fine-grained perception of single images and general comprehension across multiple images. However, existing MLLMs still face challenges in achieving precise grounding in complex multi-image scenarios. To address this, we first explore a Chain-of-Thought (CoT) framework that integrates single-image grounding with multi-image comprehension. While partially effective, it remains unstable and struggles to capture abstract visual information due to its non-end-to-end nature. Therefore, we introduce Migician, the first multi-image grounding model capable of performing free-form and accurate grounding across multiple images. To support this, we present the MGrounding-630k dataset, which comprises data for several multi-image grounding tasks derived from existing datasets, along with newly generated free-form grounding instruction-following data. Furthermore, we propose MIG-Bench, a comprehensive benchmark specifically designed for evaluating multi-image grounding capabilities. Experimental results demonstrate that our model achieves significantly superior multi-image grounding capabilities, outperforming the best existing MLLMs by 24.94% and even surpassing much larger 70B models. Our code, model, dataset, and benchmark are fully open-sourced at https://migician-vg.github.io/.

多图定位视觉语言模型数据集评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。