arXiv:2602.14788cs.CVcs.AI2026-02

通过视觉关键区域注意力提升图像指代分割精度

VIPA: Visual Informative Part Attention for Referring Image Segmentation

  • 引入视觉信息部分注意力机制,增强跨模态对齐
  • 在4个公开数据集上超越现有最佳方法
  • 适合关注细粒度图像分割与视觉语言对齐的研究者

指代图像分割(RIS)旨在根据自然语言描述分割目标物体。现有方法通过将视觉信息融入语言令牌来改进性能。为更有效利用视觉上下文以实现细粒度分割,我们提出一种新型视觉信息部分注意力(VIPA)框架。VIPA 利用视觉上下文中具有信息量的局部区域,称为视觉表达(visual expression),可为网络提供结构和语义上的目标信息,减少高方差的跨模态投影,并增强注意力机制中的语义一致性。我们还设计了一个视觉表达生成模块(VEG),该模块通过局部-全局语言上下文线索检索有信息量的视觉令牌,并精炼这些令牌以减少噪声并共享关键视觉属性。该模块使视觉表达能够综合考虑上下文,捕捉关键区域的语义视觉特征。因此,我们的框架使网络注意力能稳健地对齐细粒度关注区域。大量实验与可视化分析证明了该方法的有效性。VIPA 在四个公开的 RIS 基准测试中均优于现有最先进方法。

原文摘要 · Abstract (English)

Referring Image Segmentation (RIS) aims to segment a target object described by a natural language expression. Existing methods have evolved by leveraging the vision information into the language tokens. To more effectively exploit visual contexts for fine-grained segmentation, we propose a novel Visual Informative Part Attention (VIPA) framework for referring image segmentation. VIPA leverages the informative parts of visual contexts, called a visual expression, which can effectively provide the structural and semantic visual target information to the network. This design reduces high-variance cross-modal projection and enhances semantic consistency in an attention mechanism of the referring image segmentation. We also design a visual expression generator (VEG) module, which retrieves informative visual tokens via local-global linguistic context cues and refines the retrieved tokens for reducing noise information and sharing informative visual attributes. This module allows the visual expression to consider comprehensive contexts and capture semantic visual contexts of informative regions. In this way, our framework enables the network's attention to robustly align with the fine-grained regions of interest. Extensive experiments and visual analysis demonstrate the effectiveness of our approach. Our VIPA outperforms the existing state-of-the-art methods on four public RIS benchmarks.

图像分割视觉语言注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。