让AI理解对话中抽象意图,精准分割图像区域。
Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
- 构建对话式图像分割新任务,融合意图与物理推理。
- 提出自动生成标注数据的AI引擎,实现无监督训练。
- 模型在复杂语义分割上超越现有方法,适合多轮交互场景。
对话式图像分割将抽象、以意图为导向的概念转化为像素级掩码。现有参考图像定位研究主要关注类别和空间查询(如“最左边的苹果”),却忽略了功能与物理推理(如“我可以在哪里安全地存放刀具?”)。本文填补这一空白,提出对话式图像分割(CIS)及基准数据集ConverSeg,涵盖实体、空间关系、意图、可操作性、功能、安全性与物理推理。同时提出ConverSeg-Net,融合强分割先验与语言理解,并设计无需人工标注的AI数据生成引擎。实验表明,现有语言引导分割模型在CIS任务上表现不足,而使用该数据训练的ConverSeg-Net在ConverSeg上取得显著提升,且在已有基准上保持优异性能。
原文摘要 · Abstract (English)
Conversational image segmentation grounds abstract, intent-driven concepts into pixel-accurate masks. Prior work on referring image grounding focuses on categorical and spatial queries (e.g., "left-most apple") and overlooks functional and physical reasoning (e.g., "where can I safely store the knife?"). We address this gap and introduce Conversational Image Segmentation (CIS) and ConverSeg, a benchmark spanning entities, spatial relations, intent, affordances, functions, safety, and physical reasoning. We also present ConverSeg-Net, which fuses strong segmentation priors with language understanding, and an AI-powered data engine that generates prompt-mask pairs without human supervision. We show that current language-guided segmentation models are inadequate for CIS, while ConverSeg-Net trained on our data engine achieves significant gains on ConverSeg and maintains strong performance on existing language-guided segmentation benchmarks. Project webpage: https://glab-caltech.github.io/converseg/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。