用轻量最优传输统一多模态多视角电商表征学习
Factorized Transport Alignment for Multimodal and Multiview E-commerce Representation Learning
- 提出因子化传输框架,高效融合主图、辅图和文本视图
- 训练成本从视图数平方降为常数,推理时无额外开销
- 在百万级商品数据上提升召回率最高达7.9%,适合大模型部署
电商快速发展要求鲁棒的多模态表征以捕捉用户生成商品中丰富的信号。现有视觉-语言模型(VLMs)通常仅对齐标题与主图(单视图),忽略了在Etsy或Poshmark等开放市场中提供关键语义的非主图和辅助文本视图。为此,我们提出一种通过因子化传输(Factorized Transport)统一多模态与多视图学习的框架,这是一种轻量级最优传输近似,具备可扩展性和部署效率。训练时强调主视图并随机采样辅助视图,使训练成本从视图数的平方降至每项常数;推理时所有视图融合为单一缓存嵌入,保持双塔检索效率且无在线开销。在包含100万商品条目和30万交互的工业数据集上,该方法在跨视图和查询到商品的检索任务中持续提升性能,相较强基线最高实现Recall@500提升7.9%。整体上,该框架将最优传输学习与可扩展性结合,使大规模电商搜索的多视图预训练成为可能。
原文摘要 · Abstract (English)
The rapid growth of e-commerce requires robust multimodal representations that capture diverse signals from user-generated listings. Existing vision-language models (VLMs) typically align titles with primary images, i.e., single-view, but overlook non-primary images and auxiliary textual views that provide critical semantics in open marketplaces such as Etsy or Poshmark. To this end, we propose a framework that unifies multimodal and multi-view learning through Factorized Transport, a lightweight approximation of optimal transport, designed for scalability and deployment efficiency. During training, the method emphasizes primary views while stochastically sampling auxiliary ones, reducing training cost from quadratic in the number of views to constant per item. At inference, all views are fused into a single cached embedding, preserving the efficiency of two-tower retrieval with no additional online overhead. On an industrial dataset of 1M product listings and 0.3M interactions, our approach delivers consistent improvements in cross-view and query-to-item retrieval, achieving up to +7.9% Recall@500 over strong multimodal baselines. Overall, our framework bridges scalability with optimal transport-based learning, making multi-view pretraining practical for large-scale e-commerce search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。