arXiv:2410.18923cs.CVcs.AI2024-10被引 25

SegLLM通过多轮对话记忆实现复杂意图的交互式分割

SegLLM: Multi-round Reasoning Segmentation

  • 利用视觉与文本记忆,多轮整合分割结果进行推理
  • 在新基准MRSeg上性能领先现有方法20%以上
  • 提升单轮指代表达分割与定位准确率5.5%和4.5%

我们提出SegLLM,一种新型多轮交互式推理分割模型,通过利用视觉与文本输出的对话记忆,增强基于大模型的分割能力。该模型采用掩码感知的多模态大模型,将之前的分割结果重新融入输入流,使其能够推理复杂用户意图,并在多轮交互中基于先前识别实体的方位、互动及层级关系进行对象分割。此能力支持以对话方式响应视觉与文本查询。在新构建的MRSeg基准上评估,SegLLM在多轮交互推理分割任务中表现优于现有方法超过20%。此外,训练于多轮推理分割数据可提升标准单轮指代表达分割与定位任务性能,使指代表达分割的cIoU提升5.5%,指代表达定位的[email protected]提升4.5%。

原文摘要 · Abstract (English)

We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual outputs. By leveraging a mask-aware multimodal LLM, SegLLM re-integrates previous segmentation results into its input stream, enabling it to reason about complex user intentions and segment objects in relation to previously identified entities, including positional, interactional, and hierarchical relationships, across multiple interactions. This capability allows SegLLM to respond to visual and text queries in a chat-like manner. Evaluated on the newly curated MRSeg benchmark, SegLLM outperforms existing methods in multi-round interactive reasoning segmentation by over 20%. Additionally, we observed that training on multi-round reasoning segmentation data enhances performance on standard single-round referring segmentation and localization tasks, resulting in a 5.5% increase in cIoU for referring expression segmentation and a 4.5% improvement in [email protected] for referring expression localization.

交互分割多轮推理大模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。