arXiv:2411.16096cs.CVcs.AI2024-11被引 2

用集成与聚类优化CLIP,在数据少、图片差时提升时尚多模态搜索效果。

ENCLIP: Ensembling and Clustering-Based Contrastive Language-Image Pretraining for Fashion Multimodal Search with Limited Data and Low-Quality Images

  • 训练多个CLIP模型并集成,结合聚类增强图像表征
  • 在低质量图像和小数据集上仍保持良好检索性能
  • 适合数据稀缺、图像模糊的时尚搜索场景

多模态搜索已革新时尚产业,使用户能通过文本与图像组合发现商品。本文提出ENCLIP方法,针对时尚智能领域中数据有限和图像质量差的问题,改进对比语言-图像预训练(CLIP)模型。该方法通过训练并集成多个CLIP实例,利用聚类技术将相似图像分组,以增强表征能力。实验结果表明,该方法在低质量图像和小样本条件下仍能有效提升多模态搜索性能。ENCLIP为时尚智能领域提供了实用解决方案,显著提升了在数据稀缺环境下的模型表现。

原文摘要 · Abstract (English)

Multimodal search has revolutionized the fashion industry, providing a seamless and intuitive way for users to discover and explore fashion items. Based on their preferences, style, or specific attributes, users can search for products by combining text and image information. Text-to-image searches enable users to find visually similar items or describe products using natural language. This paper presents an innovative approach called ENCLIP, for enhancing the performance of the Contrastive Language-Image Pretraining (CLIP) model, specifically in Multimodal Search targeted towards the domain of fashion intelligence. This method focuses on addressing the challenges posed by limited data availability and low-quality images. This paper proposes an algorithm that involves training and ensembling multiple instances of the CLIP model, and leveraging clustering techniques to group similar images together. The experimental findings presented in this study provide evidence of the effectiveness of the methodology. This approach unlocks the potential of CLIP in the domain of fashion intelligence, where data scarcity and image quality issues are prevalent. Overall, the ENCLIP method represents a valuable contribution to the field of fashion intelligence and provides a practical solution for optimizing the CLIP model in scenarios with limited data and low-quality images.

多模态搜索CLIP优化时尚智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。