arXiv:2409.19569cs.CV2024-09

提出四原则对齐网络,提升图文分割精度

Fully Aligned Network for Referring Image Segmentation

论文配图:Fully Aligned Network for Referring Image Segmentation
图 1 · 摘自论文原文
  • 基于四个跨模态交互原则设计统一架构
  • 在RefCOCO等三个基准上达领先性能
  • 适合做图文理解与视觉定位的研究者

本文聚焦于指代图像分割(RIS)任务,旨在根据语言描述从图像中分割出目标物体。核心挑战在于实现模态间细粒度对齐以准确识别目标。尽管近年来基于注意力机制的跨模态交互方法已取得显著进展,但现有方法缺乏明确的交互设计准则,导致跨模态理解不足;且多数工作采用单模态掩码解码器,削弱了全模态对齐优势。为此,本文提出完全对齐网络(FAN),遵循四项跨模态交互原则,在合理规则引导下,以简洁架构实现了在主流RIS基准(RefCOCO、RefCOCO+、G-Ref)上的最优表现。

原文摘要 · Abstract (English)

This paper focuses on the Referring Image Segmentation (RIS) task, which aims to segment objects from an image based on a given language description. The critical problem of RIS is achieving fine-grained alignment between different modalities to recognize and segment the target object. Recent advances using the attention mechanism for cross-modal interaction have achieved excellent progress. However, current methods tend to lack explicit principles of interaction design as guidelines, leading to inadequate cross-modal comprehension. Additionally, most previous works use a single-modal mask decoder for prediction, losing the advantage of full cross-modal alignment. To address these challenges, we present a Fully Aligned Network (FAN) that follows four cross-modal interaction principles. Under the guidance of reasonable rules, our FAN achieves state-of-the-art performance on the prevalent RIS benchmarks (RefCOCO, RefCOCO+, G-Ref) with a simple architecture.

图像分割图文对齐跨模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。