arXiv:2603.04836cs.IR2026-03

让商品图文信息协同工作,提升电商搜索准确率

Beyond Text: Aligning Vision and Language for Multimodal E-Commerce Retrieval

  • 设计图文融合网络,统一处理文本与图像特征
  • 在大规模电商数据集上检索准确率显著提升
  • 适合做电商搜索、推荐系统优化的研究者参考

现代电商搜索本质上是多模态的:用户需综合商品文本和视觉信息做出购买决策。然而,当前工业级检索与排序系统主要依赖文本信息,未能充分利用产品图片中的丰富视觉信号。本文研究电商场景下两塔检索模型的统一文本-图像融合方法,证明领域特定微调以及查询与商品图文模态间的两阶段对齐对有效多模态检索至关重要。基于此,我们提出一种新型模态融合网络,以融合图像与文本信息并捕捉跨模态互补性。在大规模电商数据集上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Modern e-commerce search is inherently multimodal: customers make purchase decisions by jointly considering product text and visual informations. However, most industrial retrieval and ranking systems primarily rely on textual information, underutilizing the rich visual signals available in product images. In this work, we study unified text-image fusion for two-tower retrieval models in the e-commerce domain. We demonstrate that domain-specific fine-tuning and two stage alignment between query with product text and image modalities are both crucial for effective multimodal retrieval. Building on these insights, we propose a noval modality fusion network to fuse image and text information and capture cross-modal complementary information. Experiments on large-scale e-commerce datasets validate the effectiveness of the proposed approach.

多模态电商搜索图文对齐两塔模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。