用多组隐式表达增强文本描述,提升图像指代定位精度
Latent Expression Generation for Referring Image Segmentation and Grounding
- 从单个文本生成多个隐式表达,融合视觉中缺失的细节
- 在多个基准上超越现有方法,GRES任务表现最优
- 适合需要精准视觉指代的场景,如智能助手、机器人导航
视觉定位任务如指代图像分割(RIS)和指代表达理解(REC)旨在根据文本描述定位目标对象。图像中的目标可被多种方式描述,涵盖颜色、位置等多样属性。然而,现有方法多依赖单一文本输入,仅捕捉部分视觉信息,导致与相似物体混淆。为此,我们提出一种新框架,通过在单个文本输入基础上生成多个隐式表达,融入原描述中缺失的互补视觉细节。具体地,引入主体分布器与视觉概念注入模块,将共享主体与独特属性概念嵌入隐式表示,以捕捉独特且目标特定的视觉线索。同时,提出正边际对比学习策略,使所有隐式表达与原始文本对齐,同时保留细微差异。实验表明,该方法在多个基准上均优于当前最优的RIS与REC方法,并在泛化指代表达分割(GRES)基准上取得卓越性能。
原文摘要 · Abstract (English)
Visual grounding tasks, such as referring image segmentation (RIS) and referring expression comprehension (REC), aim to localize a target object based on a given textual description. The target object in an image can be described in multiple ways, reflecting diverse attributes such as color, position, and more. However, most existing methods rely on a single textual input, which captures only a fraction of the rich information available in the visual domain. This mismatch between rich visual details and sparse textual cues can lead to the misidentification of similar objects. To address this, we propose a novel visual grounding framework that leverages multiple latent expressions generated from a single textual input by incorporating complementary visual details absent from the original description. Specifically, we introduce subject distributor and visual concept injector modules to embed both shared-subject and distinct-attributes concepts into the latent representations, thereby capturing unique and target-specific visual cues. We also propose a positive-margin contrastive learning strategy to align all latent expressions with the original text while preserving subtle variations. Experimental results show that our method not only outperforms state-of-the-art RIS and REC approaches on multiple benchmarks but also achieves outstanding performance on the generalized referring expression segmentation (GRES) benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。