arXiv:2503.17871cs.CVcs.AI2025-03CVPR被引 5

用AI生成精准图像检索描述,提升复杂查询的准确性。

good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval

  • 用视觉语言模型提取物体细节并生成对比描述
  • 合成带语义变化的文本指令,提升检索精度
  • 适合做多模态检索和图像查询研究者使用

组合图像检索(CIR)允许用户通过参考图像结合文本修改来搜索图像。尽管视觉-语言模型的进步提升了CIR性能,但数据集限制仍是主要障碍。现有数据集常依赖简单、模糊或不足的手动标注,难以支持细粒度检索。我们提出good4cir,一种结构化管道,利用视觉-语言模型生成高质量合成标注。该方法包括:(1) 从查询图像中提取细粒度物体描述,(2) 为目标图像生成可比描述,(3) 合成捕捉图像间有意义变换的文本指令。该方法减少幻觉,增强修改多样性,并确保物体级一致性。应用该方法可改进现有数据集,并在多个领域创建新数据集。结果表明,在我们管道生成的数据集上训练的CIR模型检索准确率显著提升。我们已开源数据集构建框架,以支持后续CIR与多模态检索研究。

原文摘要 · Abstract (English)

Composed image retrieval (CIR) enables users to search images using a reference image combined with textual modifications. Recent advances in vision-language models have improved CIR, but dataset limitations remain a barrier. Existing datasets often rely on simplistic, ambiguous, or insufficient manual annotations, hindering fine-grained retrieval. We introduce good4cir, a structured pipeline leveraging vision-language models to generate high-quality synthetic annotations. Our method involves: (1) extracting fine-grained object descriptions from query images, (2) generating comparable descriptions for target images, and (3) synthesizing textual instructions capturing meaningful transformations between images. This reduces hallucination, enhances modification diversity, and ensures object-level consistency. Applying our method improves existing datasets and enables creating new datasets across diverse domains. Results demonstrate improved retrieval accuracy for CIR models trained on our pipeline-generated datasets. We release our dataset construction framework to support further research in CIR and multi-modal retrieval.

图像检索视觉语言模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。