提出结构化视觉表征,提升弱监督指代理解精度
Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

- 显式建模物体与关系的双层视觉表征
- 在三个数据集上达到当前最佳性能
- 适合做视觉语言对齐与弱监督学习的研究者
指代表达理解(REC)旨在定位图像中由自然语言描述的物体。在弱监督REC(WREC)中,现有方法主要基于锚点级视觉表示,即使引入辅助线索,关系信息仍隐含于单一锚点特征中,导致视觉表征扁平且仅限于单对象。本文提出结构化视觉组合表征(SVCR)框架,显式建模单对象嵌入与成对关系嵌入,构建结构化视觉空间。进一步设计组合对齐机制,在统一框架下匹配视觉与文本嵌入,实现弱监督下的组合式跨模态匹配。在RefCOCO、RefCOCO+和RefCOCOg上的实验表明,所提方法达到最优性能,验证了显式结构化表征与对齐机制的有效性。
原文摘要 · Abstract (English)
Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily operate on anchor-level visual representations. Even when enriched with auxiliary cues, relational interactions remain implicitly encoded within individual anchor features. The resulting visual representation remains flat and unary-only, limiting its ability to align with the structured nature of language. In this work, we propose a Structured Visual Compositional Representation (SVCR) learning framework for WREC. Rather than implicitly encoding relations within unary anchors, the proposed SVCR explicitly models both unary object embeddings and pairwise relational embeddings, forming a structured visual representation space. We further introduce a compositional alignment mechanism that matches unary and pairwise visual representations with their corresponding textual embeddings in a unified manner, enabling compositional visual-textual matching under weak supervision. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg show that the proposed SVCR achieves state-of-the-art performance. These results demonstrate the effectiveness of explicit structured visual representations and visual-textual alignment for WREC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。