arXiv:2507.16877cs.CVcs.AI2025-07被引 5

提出新模型与数据集,让机器更准理解多对象间的语言关系。

ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension

  • 用文本自适应机制动态推断多个目标的位置和数量
  • 在4个数据集上多实体定位与关系预测均达领先水平
  • 适合需要精准理解复杂描述的视觉语言任务研究者

指代表达理解(REC)旨在根据自然语言描述定位图像中的指定实体或区域。现有方法多聚焦单实体定位,忽略多实体场景中的复杂关系,影响准确性和可靠性。此外,缺乏带有细粒度图文关系标注的高质量数据集也制约了进展。为此,我们首先构建了名为ReMeX的关系感知多实体REC数据集,包含详细的关系与文本标注。随后提出ReMeREC框架,联合利用视觉与语言线索,在定位多个实体的同时建模其相互关系。为解决语言中隐含实体边界带来的语义模糊问题,引入文本自适应多实体感知器(TMP),通过细粒度语言线索动态推断实体数量与范围,生成区分性表示。同时,实体间关系推理模块(EIR)增强关系推理与全局场景理解能力。为进一步提升对细粒度提示的语言理解,还基于大语言模型构建了一个小规模辅助数据集EntityText。在四个基准数据集上的实验表明,ReMeREC在多实体定位与关系预测上均达到当前最优性能,显著优于现有方法。

原文摘要 · Abstract (English)

Referring Expression Comprehension (REC) aims to localize specified entities or regions in an image based on natural language descriptions. While existing methods handle single-entity localization, they often ignore complex inter-entity relationships in multi-entity scenes, limiting their accuracy and reliability. Additionally, the lack of high-quality datasets with fine-grained, paired image-text-relation annotations hinders further progress. To address this challenge, we first construct a relation-aware, multi-entity REC dataset called ReMeX, which includes detailed relationship and textual annotations. We then propose ReMeREC, a novel framework that jointly leverages visual and textual cues to localize multiple entities while modeling their inter-relations. To address the semantic ambiguity caused by implicit entity boundaries in language, we introduce the Text-adaptive Multi-entity Perceptron (TMP), which dynamically infers both the quantity and span of entities from fine-grained textual cues, producing distinctive representations. Additionally, our Entity Inter-relationship Reasoner (EIR) enhances relational reasoning and global scene understanding. To further improve language comprehension for fine-grained prompts, we also construct a small-scale auxiliary dataset, EntityText, generated using large language models. Experiments on four benchmark datasets show that ReMeREC achieves state-of-the-art performance in multi-entity grounding and relation prediction, outperforming existing approaches by a large margin.

视觉语言多实体定位关系推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。