提出双路径网络增强图像检索的上下文理解能力
HINT: Composed Image Retrieval with Dual-path Compositional Contextualized Network
- 设计双路径结构捕捉图像与文本的上下文关联
- 在两个基准数据集上所有指标均达最优表现
- 适合需要精准语义匹配的跨模态检索场景
组成式图像检索(CIR)是一项挑战性任务,旨在根据由参考图像和修改文本组成的多模态查询,在大规模图像库中检索出语义一致的目标图像。尽管现有方法在跨模态对齐和特征融合方面取得显著进展,但存在一个关键缺陷:忽视了区分匹配样本时的上下文信息。解决这一问题面临两大挑战:隐含依赖关系以及缺乏差异增强机制。为此,我们提出双路径组合式上下文网络(HINT),可实现上下文编码并放大匹配与非匹配样本间的相似度差异,从而提升复杂场景下CIR模型的性能上限。HINT在两个CIR基准数据集上所有指标均达到最优,验证了其优越性。代码已开源:https://github.com/zh-mingyu/HINT。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm. It aims to retrieve target images from large-scale image databases that are consistent with the modification semantics, based on a multimodal query composed of a reference image and modification text. Although existing methods have made significant progress in cross-modal alignment and feature fusion, a key flaw remains: the neglect of contextual information in discriminating matching samples. However, addressing this limitation is not an easy task due to two challenges: 1) implicit dependencies and 2) the lack of a differential amplification mechanism. To address these challenges, we propose a dual-patH composItional coNtextualized neTwork (HINT), which can perform contextualized encoding and amplify the similarity differences between matching and non-matching samples, thus improving the upper performance of CIR models in complex scenarios. Our HINT model achieves optimal performance on all metrics across two CIR benchmark datasets, demonstrating the superiority of our HINT model. Codes are available at https://github.com/zh-mingyu/HINT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。