arXiv:2503.19296cs.CVcs.MM2025-03被引 49

提出细粒度文本反转网络,提升零样本图像组合检索精度

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

  • 将图像拆分为主体与属性的伪词标记,更精准表达图像内容
  • 在三个基准数据集上显著优于现有方法,实现更高检索准确率
  • 适合需要零样本图像检索的视觉-语言应用开发者参考

组合图像检索(CIR)允许用户通过包含参考图像和修改文本的多模态查询搜索目标图像。然而,由于训练数据标注成本高昂,近期研究转向更具挑战性的零样本组合图像检索(ZS-CIR),即不依赖标注三元组完成任务。现有方法通过预训练文本反转网络将图像映射为单一伪词标记,将其转化为标准文本到图像检索任务。但其粗粒度表示难以准确捕捉图像全貌。为此,本文提出面向零样本组合图像检索的细粒度文本反转网络(FTI4CIR)。该方法包含两个核心组件:细粒度伪词标记映射与基于三元组描述的语义正则化。前者将图像映射为主观伪词标记和多个属性伪词标记,全面表达图像;后者基于BLIP生成的图像描述模板,联合对齐细粒度伪词标记至真实词嵌入空间。在三个基准数据集上的大量实验验证了所提方法的优越性。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) allows users to search target images with a multimodal query, comprising a reference image and a modification text that describes the user's modification demand over the reference image. Nevertheless, due to the expensive labor cost of training data annotation, recent researchers have shifted to the challenging task of zero-shot CIR (ZS-CIR), which targets fulfilling CIR without annotated triplets. The pioneer ZS-CIR studies focus on converting the CIR task into a standard text-to-image retrieval task by pre-training a textual inversion network that can map a given image into a single pseudo-word token. Despite their significant progress, their coarse-grained textual inversion may be insufficient to capture the full content of the image accurately. To overcome this issue, in this work, we propose a novel Fine-grained Textual Inversion Network for ZS-CIR, named FTI4CIR. In particular, FTI4CIR comprises two main components: fine-grained pseudo-word token mapping and tri-wise caption-based semantic regularization. The former maps the image into a subject-oriented pseudo-word token and several attribute-oriented pseudo-word tokens to comprehensively express the image in the textual form, while the latter works on jointly aligning the fine-grained pseudo-word tokens to the real-word token embedding space based on a BLIP-generated image caption template. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of our proposed method.

图像检索文本反转零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。