提出新数据集i-CIR和无训练方法BASIC,实现精准实例级图像检索。
Instance-Level Composed Image Retrieval
- 基于预训练视觉语言模型,分步计算图文-图像相似度并融合
- 在4000万干扰项中仍保持高精度,超越现有最佳方法
- 适合需要细粒度图像检索的科研与工业应用
组成图像检索(CIR)因缺乏高质量训练与评估数据而发展受限。本文提出新评估数据集i-CIR,聚焦实例级类别定义:目标是检索与视觉查询中特定物体相同的图像,且该物体需在文本查询定义的各种变化下被识别。数据集设计紧凑,通过半自动硬负样本选择,在超过4000万随机干扰项中维持挑战性。为解决高质量训练数据难题,提出无需训练的BASIC方法:分别估计查询-图像与查询-文本到图像的相似度,采用后期融合加权同时满足双条件的图像,抑制仅满足单一条件的图像。每个相似度计算通过简单直观组件进一步优化。BASIC在i-CIR上达到新基准,同时在原有语义级分类定义的CIR数据集上也表现更优。
原文摘要 · Abstract (English)
The progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge-comparable to retrieval among more than 40M random distractors-through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。