提升电商搜索中图文异构查询的匹配精度
Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search
- 设计异构模态融合模型,统一图文信息表示
- 在工业数据集上显著优于基线模型
- 适合做电商搜索与多模态匹配的研究者
语义检索通过文本查询匹配相关商品,是提升电商搜索效果的核心技术。本文研究多模态检索问题,即利用商品图像等视觉信息补充文本信息,增强商品表征以提升检索性能。尽管跨模态学习在视觉问答、媒体摘要等任务中已有广泛研究,但在查询为单模态(仅文本)而商品为多模态(图文结合)的非对称场景下,模态融合与对齐仍是未解难题。为此,本文提出SMAR模型(语义增强的模态异构检索),有效解决该场景下的模态融合与对齐问题。在工业级数据集上的大量实验表明,所提模型在检索准确率上显著优于基线方法。我们已开源该工业数据集,以支持可复现性及后续研究。
原文摘要 · Abstract (English)
Semantic retrieval, which retrieves semantically matched items given a textual query, has been an essential component to enhance system effectiveness in e-commerce search. In this paper, we study the multimodal retrieval problem, where the visual information (e.g, image) of item is leveraged as supplementary of textual information to enrich item representation and further improve retrieval performance. Though learning from cross-modality data has been studied extensively in tasks such as visual question answering or media summarization, multimodal retrieval remains a non-trivial and unsolved problem especially in the asymmetric scenario where the query is unimodal while the item is multimodal. In this paper, we propose a novel model named SMAR, which stands for Semantic-enhanced Modality-Asymmetric Retrieval, to tackle the problem of modality fusion and alignment in this kind of asymmetric scenario. Extensive experimental results on an industrial dataset show that the proposed model outperforms baseline models significantly in retrieval accuracy. We have open sourced our industrial dataset for the sake of reproducibility and future research works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。