arXiv:2507.05970cs.CV2025-07被引 2

用自动生成的三元组数据训练图像检索模型,实现零样本性能突破。

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

  • 用大语言模型生成提示词,驱动文生图模型合成带相同元素的图像对。
  • 在三个基准上实现顶尖零样本性能,首次证明全合成数据可行。
  • 适合研究视觉语言对齐、大规模数据生成与零样本检索的学者。

作为一项具有挑战性的视觉-语言任务,组合图像检索(CIR)旨在通过多模态(图像+文本)查询检索目标图像。尽管现有方法表现良好,但其依赖昂贵的人工标注三元组,限制了可扩展性与零样本能力。为此,我们提出一种可扩展的自动三元组生成流程,并构建全合成数据集CIRHS(Composed Image Retrieval on High-quality Synthetic Triplets)。该流程利用大语言模型(LLM)生成多样化提示词,控制文本到图像生成模型,产出每对图像中包含相同元素的图像对,经筛选与重组形成CIRHS数据集。此外,我们提出新型框架混合上下文对齐(CoAlign),可在更广上下文中实现全局对齐与局部推理,使模型学习更鲁棒、信息量更强的表示。使用合成数据集CIRHS,CoAlign在三个常用基准上实现卓越的零样本性能,首次验证了完全基于合成数据训练CIR模型的可行性。同时,在监督训练下,该方法超越所有现有最先进方法,证实了所提框架的有效性。代码与数据集将尽快发布。

原文摘要 · Abstract (English)

As a challenging vision-language (VL) task, Composed Image Retrieval (CIR) aims to retrieve target images using multimodal (image+text) queries. Although many existing CIR methods have attained promising performance, their reliance on costly, manually labeled triplets hinders scalability and zero-shot capability. To address this issue, we propose a scalable pipeline for automatic triplet generation, along with a fully synthetic dataset named Composed Image Retrieval on High-quality Synthetic Triplets (CIRHS). Our pipeline leverages a large language model (LLM) to generate diverse prompts, controlling a text-to-image generative model to produce image pairs with identical elements in each pair, which are then filtered and reorganized to form the CIRHS dataset. In addition, we introduce Hybrid Contextual Alignment (CoAlign), a novel CIR framework, which can accomplish global alignment and local reasoning within a broader context, enabling the model to learn more robust and informative representations. By utilizing the synthetic CIRHS dataset, CoAlign achieves outstanding zero-shot performance on three commonly used benchmarks, demonstrating for the first time the feasibility of training CIR models on a fully synthetic dataset. Furthermore, under supervised training, our method outperforms all the state-of-the-art supervised CIR approaches, validating the effectiveness of our proposed retrieval framework. The code and the CIRHS dataset will be released soon.

图像检索合成数据零样本视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。