arXiv:2503.21309cs.CVcs.AI2025-03被引 54

提出细粒度图像检索框架,精准解析文本修改意图

FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval

  • 构建细粒度标注流程,精确捕捉图像修改语义
  • 在两个数据集上实现超越现有方法的检索精度
  • 适合需要高精度图像修改检索的研究者使用

组合图像检索(CIR)通过包含参考图和修改文本的多模态查询实现图像检索:参考图定义检索上下文,修改文本指定期望变化。然而,现有CIR数据集主要使用粗粒度修改文本(CoarseMT),难以捕捉细粒度检索意图,导致正样本不准确且视觉相似图像间歧义增大,降低检索准确率,需人工筛选或重复查询。为此,我们开发了一套鲁棒的细粒度CIR数据标注流程,减少不准确正样本,提升系统对修改意图的辨识能力。基于此流程,我们重构了FashionIQ与CIRR数据集,形成Fine-FashionIQ与Fine-CIRR两个细粒度数据集。同时,提出首个显式解析修改文本的CIR框架FineCIR,能有效捕捉细粒度修改语义并对其模糊视觉实体进行对齐,显著提升检索精度。大量实验表明,FineCIR在细粒度及传统CIR基准测试中均持续优于现有先进方法。代码与数据集已开源。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) facilitates image retrieval through a multimodal query consisting of a reference image and modification text. The reference image defines the retrieval context, while the modification text specifies desired alterations. However, existing CIR datasets predominantly employ coarse-grained modification text (CoarseMT), which inadequately captures fine-grained retrieval intents. This limitation introduces two key challenges: (1) ignoring detailed differences leads to imprecise positive samples, and (2) greater ambiguity arises when retrieving visually similar images. These issues degrade retrieval accuracy, necessitating manual result filtering or repeated queries. To address these limitations, we develop a robust fine-grained CIR data annotation pipeline that minimizes imprecise positive samples and enhances CIR systems' ability to discern modification intents accurately. Using this pipeline, we refine the FashionIQ and CIRR datasets to create two fine-grained CIR datasets: Fine-FashionIQ and Fine-CIRR. Furthermore, we introduce FineCIR, the first CIR framework explicitly designed to parse the modification text. FineCIR effectively captures fine-grained modification semantics and aligns them with ambiguous visual entities, enhancing retrieval precision. Extensive experiments demonstrate that FineCIR consistently outperforms state-of-the-art CIR baselines on both fine-grained and traditional CIR benchmark datasets. Our FineCIR code and fine-grained CIR datasets are available at https://github.com/SDU-L/FineCIR.git.

图像检索细粒度理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。