arXiv:2507.04735cs.CV2025-07中稿 · Ital-IA 2025

用结构化描述提升布料图文检索准确率,发现感知编码器表现最佳。

An analysis of vision-language models for fabric retrieval

  • 用多模态大模型自动生成自然语言和结构化属性描述
  • 结构化描述使复杂布料类别的检索准确率显著提升
  • 适合工业领域跨模态检索研究者参考

跨模态检索在信息检索与推荐系统中至关重要,尤其在制造等专业领域,产品信息常以视觉样本搭配文本描述的形式存在。本文研究了视觉语言模型(VLMs)在布料样本上的零样本文本到图像检索任务。为解决公开数据集缺失问题,我们提出一种自动化标注流程,利用多模态大语言模型(MLLMs)生成两类文本描述:自由形式自然语言与结构化属性描述。在此基础上,评估了三种VLMs的表现:CLIP、LAION-CLIP和Meta的Perception Encoder。实验表明,结构化、属性丰富的描述显著提升了检索精度,尤其在视觉复杂的布料类别上;Perception Encoder因具备更强的特征对齐能力而表现最优。然而,该细粒度领域中的零样本检索仍具挑战,凸显了领域适配方法的必要性。研究强调,结合技术性文本描述与先进VLM可优化工业场景下的跨模态检索。

原文摘要 · Abstract (English)

Effective cross-modal retrieval is essential for applications like information retrieval and recommendation systems, particularly in specialized domains such as manufacturing, where product information often consists of visual samples paired with a textual description. This paper investigates the use of Vision Language Models(VLMs) for zero-shot text-to-image retrieval on fabric samples. We address the lack of publicly available datasets by introducing an automated annotation pipeline that uses Multimodal Large Language Models (MLLMs) to generate two types of textual descriptions: freeform natural language and structured attribute-based descriptions. We produce these descriptions to evaluate retrieval performance across three Vision-Language Models: CLIP, LAION-CLIP, and Meta's Perception Encoder. Our experiments demonstrate that structured, attribute-rich descriptions significantly enhance retrieval accuracy, particularly for visually complex fabric classes, with the Perception Encoder outperforming other models due to its robust feature alignment capabilities. However, zero-shot retrieval remains challenging in this fine-grained domain, underscoring the need for domain-adapted approaches. Our findings highlight the importance of combining technical textual descriptions with advanced VLMs to optimize cross-modal retrieval in industrial applications.

图文检索视觉语言模型工业应用属性描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。