arXiv:2501.07365cs.IRcs.LG2025-01中稿 · EReL@MIR WWW 2025被引 6

融合图文信息提升电商搜索相关性,效果优于纯文本检索。

Multimodal semantic retrieval for product search

  • 构建产品图文联合表征,替代仅用文本的表示方式。
  • 在购买召回率和相关性准确率上均有提升,验证多模态有效性。
  • 适合关注电商搜索优化与多模态应用的研究者或工程师。

基于文本的语义检索(又称稠密检索)在网页搜索和电商搜索中已得到广泛研究,通过比较查询与目标文档的稠密向量表示来计算相关性。产品图像在电商搜索中至关重要,是用户探索商品的关键因素,但其对语义检索的影响尚未充分研究。本文构建了电商商品的多模态表征,对比纯文本表示,探究其影响。模型在电商数据集上开发并评估,结果表明,多模态表征可提升语义检索中的购买召回率或相关性准确率。此外,通过数值分析,量化了多模态模型相较于纯文本模型所独召回的匹配项,进一步验证了多模态方案的有效性。

原文摘要 · Abstract (English)

Semantic retrieval (also known as dense retrieval) based on textual data has been extensively studied for both web search and product search application fields, where the relevance of a query and a potential target document is computed by their dense vector representation comparison. Product image is crucial for e-commerce search interactions and is a key factor for customers at product explorations. However, its impact on semantic retrieval has not been well studied yet. In this research, we build a multimodal representation for product items in e-commerce search in contrast to pure-text representation of products, and investigate the impact of such representations. The models are developed and evaluated on e-commerce datasets. We demonstrate that a multimodal representation scheme for a product can show improvement either on purchase recall or relevance accuracy in semantic retrieval. Additionally, we provide numerical analysis for exclusive matches retrieved by a multimodal semantic retrieval model versus a text-only semantic retrieval model, to demonstrate the validation of multimodal solutions.

多模态语义检索电商搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。