arXiv:2503.19910cs.CVcs.IR2025-03CVPR被引 23

用大模型自动生成图像修改数据,提升跨模态检索精度

CoLLM: A Large Language Model for Composed Image Retrieval

  • 基于图文对在线生成三元组数据,免人工标注
  • 大语言模型融合图像与修改文本,实现深度多模态理解
  • 构建340万样本数据集,显著提升检索性能

组合图像检索(CIR)旨在根据多模态查询检索图像。传统训练数据由参考图像、修改描述和目标图像组成的三元组构成,获取成本高。现有方法依赖合成三元组或利用网络爬取的图像-标题对,但前者规模小、多样性差,后者缺乏三元组结构,难以联合学习多模态查询表示。此外,现有方法在处理复杂修改文本时表现不佳。本文提出CoLLM框架,通过图像-标题对实时生成三元组,支持监督训练;利用大语言模型生成参考图像与修改文本的联合嵌入,增强多模态融合。同时构建大规模数据集MTCIR(含340万样本),并优化现有基准(CIRR和Fashion-IQ)以提升评估可靠性。实验表明,CoLLM在多个基准上达到领先性能,MTCIR使效果提升最高达15%。优化后的基准提供更可靠的评估指标,推动该领域发展。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modifications, and the target image, which are expensive and time-consuming to acquire. The scarcity of CIR datasets has led to zero-shot approaches utilizing synthetic triplets or leveraging vision-language models (VLMs) with ubiquitous web-crawled image-caption pairs. However, these methods have significant limitations: synthetic triplets suffer from limited scale, lack of diversity, and unnatural modification text, while image-caption pairs hinder joint embedding learning of the multimodal query due to the absence of triplet data. Moreover, existing approaches struggle with complex and nuanced modification texts that demand sophisticated fusion and understanding of vision and language modalities. We present CoLLM, a one-stop framework that effectively addresses these limitations. Our approach generates triplets on-the-fly from image-caption pairs, enabling supervised training without manual annotation. We leverage Large Language Models (LLMs) to generate joint embeddings of reference images and modification texts, facilitating deeper multimodal fusion. Additionally, we introduce Multi-Text CIR (MTCIR), a large-scale dataset comprising 3.4M samples, and refine existing CIR benchmarks (CIRR and Fashion-IQ) to enhance evaluation reliability. Experimental results demonstrate that CoLLM achieves state-of-the-art performance across multiple CIR benchmarks and settings. MTCIR yields competitive results, with up to 15% performance improvement. Our refined benchmarks provide more reliable evaluation metrics for CIR models, contributing to the advancement of this important field.

图像检索多模态大模型数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。