用自动生成的文本扩充数据,提升图像检索精度。
Scale Up Composed Image Retrieval Learning via Modification Text Generation
- 用大规模多模态模型生成修改文本,合成训练三元组。
- 在CIRR和FashionIQ上达到领先检索效果,召回率显著提升。
- 适合研究图像检索与数据增强的开发者参考。
组合图像检索(CIR)旨在通过参考图像与修改文本的组合查询,搜索目标图像。尽管近期取得进展,该任务仍因训练数据有限及三元组标注繁琐而面临挑战。本文提出通过合成训练三元组来扩充资源。首先利用大规模多模态模型训练修改文本生成器,在预训练阶段直接生成基于图像对的修改文本导向合成三元组(MTST)。在微调阶段,先反向生成修改文本,将目标图像回溯至参考图像,再设计两跳对齐策略逐步缩小多模态对与目标图像间的语义差距。通过循环方式学习隐式原型,再融合原型特征与修改文本,实现与目标图像的精准对齐。大量实验验证了生成三元组的有效性,所提方法在CIRR与FashionIQ基准上均取得竞争力的召回率。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) aims to search an image of interest using a combination of a reference image and modification text as the query. Despite recent advancements, this task remains challenging due to limited training data and laborious triplet annotation processes. To address this issue, this paper proposes to synthesize the training triplets to augment the training resource for the CIR problem. Specifically, we commence by training a modification text generator exploiting large-scale multimodal models and scale up the CIR learning throughout both the pretraining and fine-tuning stages. During pretraining, we leverage the trained generator to directly create Modification Text-oriented Synthetic Triplets(MTST) conditioned on pairs of images. For fine-tuning, we first synthesize reverse modification text to connect the target image back to the reference image. Subsequently, we devise a two-hop alignment strategy to incrementally close the semantic gap between the multimodal pair and the target image. We initially learn an implicit prototype utilizing both the original triplet and its reversed version in a cycle manner, followed by combining the implicit prototype feature with the modification text to facilitate accurate alignment with the target image. Extensive experiments validate the efficacy of the generated triplets and confirm that our proposed methodology attains competitive recall on both the CIRR and FashionIQ benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。