arXiv:2607.17499cs.AI2026-07

打造专为电商设计的多模态搜索模型,提升商品交易转化率。

Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation

论文配图:Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation
图 1 · 摘自论文原文
  • 构建混合语义标识机制,融合文本与图像特征。
  • 上线测试中GMV提升13.61%,交易量增8.21%。
  • 适合电商平台研发人员参考落地多模态搜索系统。

电商搜索已从简单关键词查询演变为结合商品图像、自然语言描述和混合意图指令的复杂多模态交互。现有方法面临两难:单模态专用模型独立运行,无法处理跨模态查询;通用视觉语言模型则缺乏细粒度商品理解、用户行为建模和商业意图推理所需的领域知识。本文提出Pailitao-MMSearch,一个原生电商多模态搜索基础模型,引入三项创新:(1) 混合语义标识(HybSID);(2) 两阶段持续预训练策略;(3) 混合推理后训练流程。基于Qwen架构,在淘宝“Pailitao”多模态搜索平台部署,线上A/B测试显示,相比传统多模态搜索管道,最高实现GMV提升13.61%、交易量增长8.21%,验证了原生电商多模态大模型的有效性。

原文摘要 · Abstract (English)

The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that seamlessly combine product images, natural language descriptions, and mixed-intent instructions. However, existing approaches face a critical dilemma: single-modal specialist models, deployed independently for text retrieval, visual search, and voice recognition, operate in isolation and cannot handle cross-modal queries, while general-purpose vision-language models lack the domain-specific knowledge necessary for fine-grained product understanding, user behavior modeling, and commercial intent reasoning. In this work, we present Pailitao-MMSearch, one native e-commerce multimodal search foundation model designed to bridge this gap. Our approach introduces three key innovations: (1)HybSID (Hybrid Semantic ID);(2)a two-stage continual pre-training strategy; and (3)a hybrid reasoning post-training pipeline. Built upon Qwen and deployed on Taobao's Pailitao multimodal search platform, Pailitao-MMSearch achieves substantial improvements in online A/B testing, including up to +13.61\% in Gross Merchandise Volume (GMV) and +8.21\% in transaction volume compared to traditional multi-modal search pipeline, demonstrating the effectiveness of our native e-commerce multimodal search large language models.

多模态搜索电商大模型视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。