arXiv:2512.06862cs.CV2025-12被引 3

支持文本和多种视觉提示的通用图像分割新任务

Omni-Referring Image Segmentation

  • 融合文本与多种视觉提示进行图像分割
  • 构建包含18.7万条提示的大规模数据集OmniRef
  • 适合需要多模态交互的通用分割场景

本文提出一项新型任务——全模态指代图像分割(Omni-Referring Image Segmentation, OmniRIS),旨在实现高度泛化的图像分割。相较于现有单模态条件分割任务(如RIS和视觉RIS),OmniRIS支持文本指令以及带有掩码、边界框或草图的参考图像作为全模态提示,从而充分结合文本的细粒度属性指代和视觉模态对罕见物体的精准定位优势。此外,OmniRIS可处理一对一、多对多等多种分割场景,提升实际应用灵活性。为推动该任务研究,本文构建了大规模数据集OmniRef,包含30,956张图像上的186,939条全模态提示,并建立全面评估体系。同时提出强而通用的基线模型OmniSegNet,以应对全模态提示编码等关键挑战。大量实验不仅验证了OmniSegNet在遵循全模态指令方面的能力,也展示了OmniRIS在高度泛化图像分割中的优越性。

原文摘要 · Abstract (English)

In this paper, we propose a novel task termed Omni-Referring Image Segmentation (OmniRIS) towards highly generalized image segmentation. Compared with existing unimodally conditioned segmentation tasks, such as RIS and visual RIS, OmniRIS supports the input of text instructions and reference images with masks, boxes or scribbles as omni-prompts. This property makes it can well exploit the intrinsic merits of both text and visual modalities, i.e., granular attribute referring and uncommon object grounding, respectively. Besides, OmniRIS can also handle various segmentation settings, such as one v.s. many and many v.s. many, further facilitating its practical use. To promote the research of OmniRIS, we also rigorously design and construct a large dataset termed OmniRef, which consists of 186,939 omni-prompts for 30,956 images, and establish a comprehensive evaluation system. Moreover, a strong and general baseline termed OmniSegNet is also proposed to tackle the key challenges of OmniRIS, such as omni-prompt encoding. The extensive experiments not only validate the capability of OmniSegNet in following omni-modal instructions, but also show the superiority of OmniRIS for highly generalized image segmentation.

图像分割多模态指代理解通用分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。