arXiv:2602.11733cs.CVcs.AI2026-02Conference of the …

让通用视觉语言模型更好理解电商图文信息。

Adapting Vision-Language Models for E-commerce Understanding at Scale

  • 针对电商特点微调通用视觉语言模型
  • 在保持通用能力前提下显著提升电商理解性能
  • 适合需要高精度商品理解的工业场景

电商产品理解天然要求对文本、图像和结构化属性进行强多模态理解。通用视觉语言模型(VLMs)能实现可泛化的多模态表征,但目前尚无已知且公认的方法,可在不牺牲通用性能的前提下,适配电商数据中以属性为中心、多图输入及噪声较多的特点。本文通过大规模实验研究,展示了针对性地微调通用VLMs,可显著提升其在电商场景下的表现,同时保留广泛的多模态能力。此外,我们提出一个全新的综合性评估体系,涵盖深层商品理解、严格指令遵循与动态属性提取三个维度。

原文摘要 · Abstract (English)

E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction.

视觉语言模型电商理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。