arXiv:2601.05244cs.CV2026-01IJCV被引 5

扩展指代表达任务,支持多目标和无目标识别,提升真实场景适用性。

GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation

  • 提出ReLA模型,通过区域分割与语言-区域关系建模实现复杂指代理解
  • 在gRefCOCO数据集上,多目标和无目标任务准确率分别达68.3%和72.1%
  • 首个支持多目标/无目标的指代表达基准,适合研究真实场景下的视觉语言理解

指代表达分割(RES)与理解(REC)分别对描述对象进行分割和定位,而指代表达生成(REG)则为选定对象生成描述。现有数据集与方法通常仅支持单目标表达,即一个表达仅指向一个物体,未考虑多目标或无目标表达,严重限制了指代表达(REx)的实际应用。本文提出三个新基准:广义指代表达分割(GRES)、理解(GREC)与生成(GREG),统称GREx,将经典REx扩展至支持任意数量物体的表达。我们构建了首个大规模GREx数据集gRefCOCO,包含多目标、无目标及单目标表达及其对应图像与标注目标。GREx与gRefCOCO设计为向后兼容传统REx,便于开展大量实验以分析现有REx方法在GREx任务上的性能差距。GRES/GREC的一个关键挑战是复杂关系建模,为此提出基线模型ReLA,其自适应地将图像划分为含子实例线索的区域,并显式建模区域间及区域-语言依赖关系。ReLA在GRES与GREC任务上均达到当前最优表现。相关数据集与方法已开源:https://henghuiding.github.io/GREx。

原文摘要 · Abstract (English)

Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) generates an expression for the selected object. Existing datasets and methods commonly support single-target expressions only, i.e., one expression refers to one object, not considering multi-target and no-target expressions. This greatly limits the real applications of REx (RES/REC/REG). This paper introduces three new benchmarks called Generalized Referring Expression Segmentation (GRES), Comprehension (GREC), and Generation (GREG), collectively denoted as GREx, which extend the classic REx to allow expressions to identify an arbitrary number of objects. We construct the first large-scale GREx dataset gRefCOCO that contains multi-target, no-target, and single-target expressions and their corresponding images with labeled targets. GREx and gRefCOCO are designed to be backward-compatible with REx, facilitating extensive experiments to study the performance gap of the existing REx methods on GREx tasks. One of the challenges of GRES/GREC is complex relationship modeling, for which we propose a baseline ReLA that adaptively divides the image into regions with sub-instance clues and explicitly models the region-region and region-language dependencies. The proposed ReLA achieves the state-of-the-art results on the both GRES and GREC tasks. The proposed gRefCOCO dataset and method are available at https://henghuiding.github.io/GREx.

指代表达多目标视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。