arXiv:2606.19684cs.CV2026-06

用多模态大模型生成属性感知三元组,提升服装图像细粒度检索效果

Exploring Multi-Modal Large Language Models and Two-Stage Fine-Tuning for Fashion Image Retrieval

论文配图:Exploring Multi-Modal Large Language Models and Two-Stage Fine-Tuning for Fashion Image Retrieval
图 1 · 摘自论文原文
  • 引入LLaVA生成带属性的三元组,增强对颜色纹理等细节的理解
  • 两阶段微调使对比学习更有效,准确率显著提升
  • 适合需要精细属性匹配的服装检索场景

组合图像检索通过参考图和修改后的文本描述来查找目标图像。在时尚领域,该任务需理解颜色、图案、纹理等细微属性变化。然而,现有方法受限于标注数据稀少和简单的负样本采样。我们提出一个新框架,整合多模态大语言模型(LLaVA)生成属性感知三元组,并引入两阶段微调策略以增强对比学习。利用预训练视觉语言模型(如CLIP-ViT/B32),将句级提示与相对描述拼接,通过静态表示扩大负样本数量。实验表明,该方法提升了组合推理能力与细粒度检索表现,验证了框架在时尚检索中的可行性和潜力。

原文摘要 · Abstract (English)

Composed image retrieval retrieves a target image using a composed query of a reference image and a modified text description. In the fashion domain, this task requires understanding subtle attribute variations such as color, pattern, and texture. However, existing approaches face limitations due to scarce annotated data and simplistic negative sampling. We propose a novel framework that integrates a multi-modal large language model (LLaVA) to generate attribute-aware triplets and introduces a two-stage fine-tuning strategy to enhance contrastive learning. We leverage pretrained vision-language models, such as CLIP-ViT/B32, to generate and concatenate sentence-level prompts with the relative caption and to scale the number of negatives using static representations. Experimental results demonstrate enhanced compositional reasoning and improved fine-grained retrieval behavior, underscoring the feasibility and potential of the proposed framework for fashion retrieval.

图像检索多模态大模型时尚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。