arXiv:2507.07135cs.LG2025-07被引 2

构建大规模时尚细粒度图像检索数据集,提升电商场景下精准修改指令的图像搜索能力。

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval

  • 用视觉语言模型与大模型自动构建时尚图像修改文本,降低人工标注成本。
  • 在Fashion IQ和enhFashionIQ上,细粒度检索准确率显著提升,尤其对复杂指令有效。
  • 适合做时尚电商、个性化推荐或细粒度视觉语言理解的研究者使用。

组合图像检索(CIR)任务是根据参考图像和修改文本检索目标图像。现有方法依赖大型预训练视觉-语言模型(VLM),在颜色、纹理等通用概念上表现良好,但在时尚领域仍面临挑战,因时尚语义丰富多样,需精细的视觉与语言理解能力。此外,受限于专业标注成本,缺乏大规模带详细标注的时尚数据集。为此,我们提出FACap,一个大规模、自动生成的时尚领域CIR数据集。该数据集利用网络获取的时尚图像,并通过由VLM和大语言模型(LLM)驱动的两阶段标注流程生成准确且详尽的修改文本。随后,我们提出新模型FashionBLIP-2,其在FACap上微调通用域的BLIP-2模型,采用轻量级适配器和多头查询-候选匹配机制,以更好地捕捉细粒度时尚信息。在Fashion IQ基准和增强版评估集enhFashionIQ上,经我们管道生成更高质量标注后进行评估,结果表明,FashionBLIP-2结合FACap预训练能显著提升模型在时尚CIR中的性能,尤其在细粒度修改文本下的检索表现突出,验证了该数据集与方法在电商等高要求场景中的价值。代码已公开。

原文摘要 · Abstract (English)

The composed image retrieval (CIR) task is to retrieve target images given a reference image and a modification text. Recent methods for CIR leverage large pretrained vision-language models (VLMs) and achieve good performance on general-domain concepts like color and texture. However, they still struggle with application domains like fashion, because the rich and diverse vocabulary used in fashion requires specific fine-grained vision and language understanding. An additional difficulty is the lack of large-scale fashion datasets with detailed and relevant annotations, due to the expensive cost of manual annotation by specialists. To address these challenges, we introduce FACap, a large-scale, automatically constructed fashion-domain CIR dataset. It leverages web-sourced fashion images and a two-stage annotation pipeline powered by a VLM and a large language model (LLM) to generate accurate and detailed modification texts. Then, we propose a new CIR model FashionBLIP-2, which fine-tunes the general-domain BLIP-2 model on FACap with lightweight adapters and multi-head query-candidate matching to better account for fine-grained fashion-specific information. FashionBLIP-2 is evaluated with and without additional fine-tuning on the Fashion IQ benchmark and the enhanced evaluation dataset enhFashionIQ, leveraging our pipeline to obtain higher-quality annotations. Experimental results show that the combination of FashionBLIP-2 and pretraining with FACap significantly improves the model's performance in fashion CIR especially for retrieval with fine-grained modification texts, demonstrating the value of our dataset and approach in a highly demanding environment such as e-commerce websites. Code is available at https://fgxaos.github.io/facap-paper-website/.

图像检索时尚电商细粒度理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。