arXiv:2409.13407cs.CV2024-09被引 8

让大模型按指令自动调整图像分割精细度,实现从整体到细节的灵活理解。

Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model

  • 根据用户指令动态调整分割粒度,支持从整体到细粒度的多级分割与描述。
  • 构建包含10,000张图像、30,000+问答对的多粒度基准数据集,填补领域空白。
  • 提出统一数据格式,提升多任务学习中视觉与语义关联能力,适合多模态研究者使用。

大型多模态模型(LMMs)在扩展大语言模型的基础上取得显著进展,最新发展表明其可通过整合分割模型生成稠密像素级分割结果。然而,现有方法的文本响应与分割掩码仍局限于实例级别,即便给出详细文本提示,也难以实现细粒度理解与分割。为克服这一局限,我们提出多粒度大型多模态模型(MGLMM),可无缝响应用户指令,将分割与描述(SegCap)粒度从全景式调整至细粒度。我们定义了新任务:多粒度分割与描述(MGSC)。由于缺乏对应基准,我们基于自研自动化标注流程建立了首个对齐掩码与描述的多粒度基准,包含10,000张图像和超过30,000个图像-问题对,并将公开数据集与标注工具以供后续研究。此外,我们提出一种新型统一的SegCap数据格式,有效融合异构分割数据集,在多任务训练中促进对象概念与视觉特征的关联。大量实验表明,我们的MGLMM在超过八项下游任务中表现卓越,涵盖MGSC、GCG、图像描述、指代表达分割、多重与空分割、推理分割等,达到当前最优性能,展现出强大潜力与广泛应用前景。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have achieved significant progress by extending large language models. Building on this progress, the latest developments in LMMs demonstrate the ability to generate dense pixel-wise segmentation through the integration of segmentation models.Despite the innovations, the textual responses and segmentation masks of existing works remain at the instance level, showing limited ability to perform fine-grained understanding and segmentation even provided with detailed textual cues.To overcome this limitation, we introduce a Multi-Granularity Large Multimodal Model (MGLMM), which is capable of seamlessly adjusting the granularity of Segmentation and Captioning (SegCap) following user instructions, from panoptic SegCap to fine-grained SegCap. We name such a new task Multi-Granularity Segmentation and Captioning (MGSC). Observing the lack of a benchmark for model training and evaluation over the MGSC task, we establish a benchmark with aligned masks and captions in multi-granularity using our customized automated annotation pipeline. This benchmark comprises 10K images and more than 30K image-question pairs. We will release our dataset along with the implementation of our automated dataset annotation pipeline for further research.Besides, we propose a novel unified SegCap data format to unify heterogeneous segmentation datasets; it effectively facilitates learning to associate object concepts with visual features during multi-task training. Extensive experiments demonstrate that our MGLMM excels at tackling more than eight downstream tasks and achieves state-of-the-art performance in MGSC, GCG, image captioning, referring segmentation, multiple and empty segmentation, and reasoning segmentation tasks. The great performance and versatility of MGLMM underscore its potential impact on advancing multimodal research.

多粒度分割多模态模型图像描述统一数据格式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。