让通用视觉语言模型更好理解电商图文信息。
Adapting Vision-Language Models for E-commerce Understanding at Scale
- 针对电商特点微调通用视觉语言模型
- 在保持通用能力前提下显著提升电商理解性能
- 适合需要高精度商品理解的工业场景
电商产品理解天然要求对文本、图像和结构化属性进行强多模态理解。通用视觉语言模型(VLMs)能实现可泛化的多模态表征,但目前尚无已知且公认的方法,可在不牺牲通用性能的前提下,适配电商数据中以属性为中心、多图输入及噪声较多的特点。本文通过大规模实验研究,展示了针对性地微调通用VLMs,可显著提升其在电商场景下的表现,同时保留广泛的多模态能力。此外,我们提出一个全新的综合性评估体系,涵盖深层商品理解、严格指令遵循与动态属性提取三个维度。
原文摘要 · Abstract (English)
E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。