对比GPT-4o-mini与Gemini 2.0 Flash在零样本下识别服装细粒度属性的能力。
Can GPT-4o mini and Gemini 2.0 Flash Predict Fine-Grained Fashion Product Attributes? A Zero-Shot Analysis
- 仅用图像作为输入,零样本评估两大模型对18类服装属性的识别能力。
- Gemini 2.0 Flash宏平均F1达56.79%,优于GPT-4o-mini的43.28%。
- 揭示模型短板,为电商场景部署和领域微调提供实用指导。
时尚零售的核心在于产品理解能力。产品属性识别有助于根据业务流程理解商品,提升客户在海量商品中浏览的体验,实现更有序的商品目录,直接影响客户的‘发现体验’。尽管大语言模型(LLMs)在多模态数据理解上表现卓越,其在细粒度时尚属性识别上的表现仍缺乏研究。本文对性能与速度、成本效率平衡的前沿模型——GPT-4o-mini与Gemini 2.0 Flash进行零样本评估,使用DeepFashion-MultiModal数据集,在18类时尚属性上展开分析。仅以图像为唯一输入,构建受限环境。结果表明,Gemini 2.0 Flash整体表现最优,所有属性的宏平均F1为56.79%,而GPT-4o-mini为43.28%。通过详细错误分析,研究为生产环境中部署这些模型于电商产品属性任务提供了实践洞察,并强调了领域特定微调的必要性。本工作也为未来时尚AI与多模态属性提取研究奠定基础。
原文摘要 · Abstract (English)
The fashion retail business is centered around the capacity to comprehend products. Product attribution helps in comprehending products depending on the business process. Quality attribution improves the customer experience as they navigate through millions of products offered by a retail website. It leads to well-organized product catalogs. In the end, product attribution directly impacts the 'discovery experience' of the customer. Although large language models (LLMs) have shown remarkable capabilities in understanding multimodal data, their performance on fine-grained fashion attribute recognition remains under-explored. This paper presents a zero-shot evaluation of state-of-the-art LLMs that balance performance with speed and cost efficiency, mainly GPT-4o-mini and Gemini 2.0 Flash. We have used the dataset DeepFashion-MultiModal (https://github.com/yumingj/DeepFashion-MultiModal) to evaluate these models in the attribution tasks of fashion products. Our study evaluates these models across 18 categories of fashion attributes, offering insight into where these models excel. We only use images as the sole input for product information to create a constrained environment. Our analysis shows that Gemini 2.0 Flash demonstrates the strongest overall performance with a macro F1 score of 56.79% across all attributes, while GPT-4o-mini scored a macro F1 score of 43.28%. Through detailed error analysis, our findings provide practical insights for deploying these LLMs in production e-commerce product attribution-related tasks and highlight the need for domain-specific fine-tuning approaches. This work also lays the groundwork for future research in fashion AI and multimodal attribute extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。