arXiv:2505.20764cs.CVcs.LG2025-05CVPR被引 9

让文字描述更精准匹配图像修改,提升组合图像检索效果。

ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval

  • 用文本中的名词短语引导图像注意力,对齐语义变化区域。
  • 在CIRR和CIRCO上达成新最好结果,零样本设置也显著领先。
  • 自动生成合成数据训练,适配无标注图像场景。

组合图像检索(CIR)旨在根据查询图像和描述其语义修改的相对文本,检索目标图像。现有方法难以准确建模图像与文本修改之间的关系,导致性能不足。为此,我们提出ConText-CIR框架,引入文本概念一致性损失,促使文本中名词短语的表示更关注查询图像的相关区域。为支持该损失的训练,我们设计了一种从现有CIR数据集或未标注图像生成合成数据的流水线。实验表明,这些组件协同工作显著提升了CIR性能,在多个基准数据集(包括CIRR和CIRCO)的监督与零样本设置下均达到新最佳水平。源代码、模型权重及新数据集已开源。

原文摘要 · Abstract (English)

Composed image retrieval (CIR) is the task of retrieving a target image specified by a query image and a relative text that describes a semantic modification to the query image. Existing methods in CIR struggle to accurately represent the image and the text modification, resulting in subpar performance. To address this limitation, we introduce a CIR framework, ConText-CIR, trained with a Text Concept-Consistency loss that encourages the representations of noun phrases in the text modification to better attend to the relevant parts of the query image. To support training with this loss function, we also propose a synthetic data generation pipeline that creates training data from existing CIR datasets or unlabeled images. We show that these components together enable stronger performance on CIR tasks, setting a new state-of-the-art in composed image retrieval in both the supervised and zero-shot settings on multiple benchmark datasets, including CIRR and CIRCO. Source code, model checkpoints, and our new datasets are available at https://github.com/mvrl/ConText-CIR.

图像检索文本对齐合成数据零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。