提出新模型解决复杂指代表达理解,支持单目标、多目标和无目标场景。
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
- 分层对齐机制融合词-对象、短语-对象与文本-图像三层次交互
- 自适应计数模块动态确定目标数量,准确率提升显著
- 适用于多种指代表达任务,通用性强,适合实际应用
本文针对具有挑战性的广义指代表达理解(GREC)任务。相较于经典仅处理单目标的指代表达理解(REC),GREC涵盖无目标和多目标表达,更贴近实际应用。现有REC方法在处理GREC复杂情况时受限于固定输出和多模态表示能力不足。为此,本文提出分层对齐增强的自适应定位网络(HieA2G),可灵活应对各类指代表达。首先,设计分层多模态语义对齐(HMSA)模块,实现词-对象、短语-对象、文本-图像三个层级的跨模态交互,提升复杂场景下的定位能力。其次,引入自适应定位计数器(AGC)动态预测目标数量,并通过辅助对比损失强化计数能力,使相同数量的对象特征拉近,不同数量则推开。大量实验表明,HieA2G在挑战性的GREC任务及其他四个任务(REC、短语定位、指代表达分割RES、广义指代表达分割GRES)上均达到最新水平,展现卓越性能与强泛化能力。
原文摘要 · Abstract (English)
In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC extends the scope to a more practical setting by further encompassing no-target and multi-target expressions. Existing REC methods face challenges in handling the complex cases encountered in GREC, primarily due to their fixed output and limitations in multi-modal representations. To address these issues, we propose a Hierarchical Alignment-enhanced Adaptive Grounding Network (HieA2G) for GREC, which can flexibly deal with various types of referring expressions. First, a Hierarchical Multi-modal Semantic Alignment (HMSA) module is proposed to incorporate three levels of alignments, including word-object, phrase-object, and text-image alignment. It enables hierarchical cross-modal interactions across multiple levels to achieve comprehensive and robust multi-modal understanding, greatly enhancing grounding ability for complex cases. Then, to address the varying number of target objects in GREC, we introduce an Adaptive Grounding Counter (AGC) to dynamically determine the number of output targets. Additionally, an auxiliary contrastive loss is employed in AGC to enhance object-counting ability by pulling in multi-modal features with the same counting and pushing away those with different counting. Extensive experimental results show that HieA2G achieves new state-of-the-art performance on the challenging GREC task and also the other 4 tasks, including REC, Phrase Grounding, Referring Expression Segmentation (RES), and Generalized Referring Expression Segmentation (GRES), demonstrating the remarkable superiority and generalizability of the proposed HieA2G.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。