arXiv:2409.03385cs.CVcs.MM2024-09被引 2

用表达式引导的动态门控和回归,让图模型重振旗鼓。

Make Graph-based Referring Expression Comprehension Great Again through Expression-guided Dynamic Gating and Regression

  • 通过子表达式引导动态门控,自动屏蔽无关目标及其连接
  • 引入表达式引导回归,显著提升定位精度
  • 无需预训练,性能超越当前最优的变压器方法

主流观点认为,基于Transformer的方法在大规模数据预训练后,已显著优于传统图模型。我们发现,多数图方法依赖通用检测器生成候选对象,面临两大挑战:一是推理时大量无关物体带来噪声,二是检测器本身定位不准。为此,提出由子表达式引导的动态门控约束(DGC)模块,可自适应地禁用无关候选及图中连接;同时设计表达式引导回归策略(EGR),用于精炼定位结果。在RefCOCO、RefCOCO+、RefCOCOg、Flickr30K、RefClef和Ref-reasoning等6个数据集上的实验表明,DGC与EGR能持续提升多种图模型性能。无需预训练,所提方法即达到甚至超越当前最先进的基于Transformer的方法。

原文摘要 · Abstract (English)

One common belief is that with complex models and pre-training on large-scale datasets, transformer-based methods for referring expression comprehension (REC) perform much better than existing graph-based methods. We observe that since most graph-based methods adopt an off-the-shelf detector to locate candidate objects (i.e., regions detected by the object detector), they face two challenges that result in subpar performance: (1) the presence of significant noise caused by numerous irrelevant objects during reasoning, and (2) inaccurate localization outcomes attributed to the provided detector. To address these issues, we introduce a plug-and-adapt module guided by sub-expressions, called dynamic gate constraint (DGC), which can adaptively disable irrelevant proposals and their connections in graphs during reasoning. We further introduce an expression-guided regression strategy (EGR) to refine location prediction. Extensive experimental results on the RefCOCO, RefCOCO+, RefCOCOg, Flickr30K, RefClef, and Ref-reasoning datasets demonstrate the effectiveness of the DGC module and the EGR strategy in consistently boosting the performances of various graph-based REC methods. Without any pretaining, the proposed graph-based method achieves better performance than the state-of-the-art (SOTA) transformer-based methods.

图模型视觉理解表达式理解定位优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。