arXiv:2503.05204cs.CV2025-03被引 5

提出轻量框架,让零样本图像检索更高效准确

Data-Efficient Generalization for Zero-shot Composed Image Retrieval

  • 引入文本补充与语义集合模块,缓解模态差异
  • 仅用少量数据即超越现有最佳方法性能
  • 适合资源受限场景下的零样本检索应用

零样本组合图像检索(ZS-CIR)旨在仅凭参考图像和文本描述,从无训练三元组的分布中检索目标图像。主流方法采用视觉-语言预训练范式,通过映射网络将图像嵌入转换为文本嵌入空间中的伪词标记。然而,该方法因模态差异及训练与推理间分布偏移,易导致模型泛化能力下降。为此,我们提出数据高效的泛化(DeG)框架,包含两项新设计:文本补充(TS)模块在训练中利用组合文本语义,增强伪词标记的语言语义,有效缓解模态差异;语义集(S-Set)利用预训练视觉-语言模型(VLMs)的零样本能力,减轻分布偏移并缓解大规模图文数据冗余带来的过拟合问题。在四个ZS-CIR基准上的大量实验表明,DeG以远少于现有方法的训练数据实现更优性能,并显著降低训练与推理时间,适用于实际部署。

原文摘要 · Abstract (English)

Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve the target image based on a reference image and a text description without requiring in-distribution triplets for training. One prevalent approach follows the vision-language pretraining paradigm that employs a mapping network to transfer the image embedding to a pseudo-word token in the text embedding space. However, this approach tends to impede network generalization due to modality discrepancy and distribution shift between training and inference. To this end, we propose a Data-efficient Generalization (DeG) framework, including two novel designs, namely, Textual Supplement (TS) module and Semantic-Set (S-Set). The TS module exploits compositional textual semantics during training, enhancing the pseudo-word token with more linguistic semantics and thus mitigating the modality discrepancy effectively. The S-Set exploits the zero-shot capability of pretrained Vision-Language Models (VLMs), alleviating the distribution shift and mitigating the overfitting issue from the redundancy of the large-scale image-text data. Extensive experiments over four ZS-CIR benchmarks show that DeG outperforms the state-of-the-art (SOTA) methods with much less training data, and saves substantial training and inference time for practical usage.

零样本检索视觉语言模型数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。