arXiv:2502.09089cs.IR2025-02被引 2

用多领域语言模型提升沃尔玛广告检索准确率

Semantic Ads Retrieval at Walmart eCommerce with Language Models Progressively Trained on Multiple Knowledge Domains

  • 用商品类别预训练BERT模型,增强对商品语义理解
  • 采用双塔网络结构,提升训练效率,搜索相关性提升16%
  • 引入人机协同渐进融合训练,适合电商广告系统优化

电商平台的赞助搜索面临诸多独特而复杂的挑战,包括搜索词与商品名称之间语言结构不对称、用户搜索意图固有的模糊性,以及海量稀疏且不平衡的搜索语料数据。检索模块在赞助搜索系统中至关重要,直接影响后续排序与竞价系统。本文提出一个端到端解决方案,优化沃尔玛.com的广告检索系统。首先,利用商品类别信息对类BERT分类模型进行预训练,增强模型对沃尔玛商品语义的理解;其次,设计双塔孪生网络结构以优化嵌入表示,提升训练效率;第三,引入人机协同渐进融合训练方法,确保模型鲁棒性。实验结果表明,该流程可使搜索相关性指标相比基线DSSM模型最高提升16%。此外,大规模在线A/B测试显示,该方法广告收入超越现有生产模型。

原文摘要 · Abstract (English)

Sponsored search in e-commerce poses several unique and complex challenges. These challenges stem from factors such as the asymmetric language structure between search queries and product names, the inherent ambiguity in user search intent, and the vast volume of sparse and imbalanced search corpus data. The role of the retrieval component within a sponsored search system is pivotal, serving as the initial step that directly affects the subsequent ranking and bidding systems. In this paper, we present an end-to-end solution tailored to optimize the ads retrieval system on Walmart.com. Our approach is to pretrain the BERT-like classification model with product category information, enhancing the model's understanding of Walmart product semantics. Second, we design a two-tower Siamese Network structure for embedding structures to augment training efficiency. Third, we introduce a Human-in-the-loop Progressive Fusion Training method to ensure robust model performance. Our results demonstrate the effectiveness of this pipeline. It enhances the search relevance metric by up to 16% compared to a baseline DSSM-based model. Moreover, our large-scale online A/B testing demonstrates that our approach surpasses the ad revenue of the existing production model.

广告检索语言模型电商推荐双塔网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。