arXiv:2503.13881cs.CV2025-03ICLR被引 35

构建大规模多目标细粒度推理分割数据集,提升模型理解复杂指令能力。

MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation

  • 构建包含19.4万条指令的多目标多粒度数据集,支持对象与部件级理解。
  • 在多目标多粒度场景下,新方法显著优于现有模型,实现有效推理。
  • 适合研究视觉语言交互、机器人任务规划及细粒度语义分割的学者使用。

大语言模型与视觉模型的融合为用户交互式视觉语言任务开辟了新可能。其中,推理分割要求模型通过理解人类指令中的隐含含义生成像素级分割掩码。然而,顺畅的人机交互不仅需要识别对象,还需理解对象及其组成部分的功能,尤其在多目标场景中更为关键。例如,当指令为“打开电视”时,可能涉及电视本身或遥控器等多种操作对象(多目标),并需理解其按钮等部件(部件级)。当前推理分割数据集多聚焦单一目标对象级推理,难以支持多目标场景下的细粒度识别。为此,我们构建了大规模数据集MMR,包含19.4万条考虑多目标、对象级与部件级的复杂隐含指令,基于已有图像-掩码集构建。该数据集通过层级化提供对象与部件信息,支持多样且上下文感知的交互。此外,我们提出一种简单而有效的多目标、对象级与部件级推理分割框架。在MMR上的实验表明,所提方法能在多目标多粒度场景中实现有效推理,而现有模型仍有提升空间。

原文摘要 · Abstract (English)

The fusion of Large Language Models with vision models is pioneering new possibilities in user-interactive vision-language tasks. A notable application is reasoning segmentation, where models generate pixel-level segmentation masks by comprehending implicit meanings in human instructions. However, seamless human-AI interaction demands more than just object-level recognition; it requires understanding both objects and the functions of their detailed parts, particularly in multi-target scenarios. For example, when instructing a robot to \textit{turn on the TV"}, there could be various ways to accomplish this command. Recognizing multiple objects capable of turning on the TV, such as the TV itself or a remote control (multi-target), provides more flexible options and aids in finding the optimized scenario. Furthermore, understanding specific parts of these objects, like the TV's button or the remote's button (part-level), is important for completing the action. Unfortunately, current reasoning segmentation datasets predominantly focus on a single target object-level reasoning, which limits the detailed recognition of an object's parts in multi-target contexts. To address this gap, we construct a large-scale dataset called Multi-target and Multi-granularity Reasoning (MMR). MMR comprises 194K complex and implicit instructions that consider multi-target, object-level, and part-level aspects, based on pre-existing image-mask sets. This dataset supports diverse and context-aware interactions by hierarchically providing object and part information. Moreover, we propose a straightforward yet effective framework for multi-target, object-level, and part-level reasoning segmentation. Experimental results on MMR show that the proposed method can reason effectively in multi-target and multi-granularity scenarios, while the existing reasoning segmentation model still has room for improvement.

推理分割多目标细粒度视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。