arXiv:2511.00694cs.IR2025-11中稿 · 2025 IEEE Internat…

基于商品分类的负样本挖掘,提升电商搜索个性化召回效果。

Taxonomy-based Negative Sampling In Personalized Semantic Search for E-commerce

  • 用商品分类构建困难负样本,增强语义区分能力。
  • 离线测试召回率优于传统方法,线上转化率显著提升。
  • 减少训练开销,适合大规模电商系统部署。

大型零售平台商品具有领域特性,需模型理解相似商品间的细微差异。现有训练采样方法往往计算成本高或难以实施,且未考虑用户历史购买行为,导致检索结果不相关。本文提出一种电商搜索语义检索模型,将查询与商品映射至共享向量空间,并引入新型基于分类的困难负样本采样(TB-HNS)策略,挖掘上下文相关但具有挑战性的负样本。为实现个性化检索,模型还融合用户级偏好,建模其过往购买历史与行为。离线实验表明,该方法在Recall@K上优于BM25、ANCE及主流神经基线;线上A/B测试显示,转化率、加购率和客单价均有显著提升。同时验证了分类驱动负样本能降低训练开销并加速收敛,并分享了系统规模化落地的实践经验。

原文摘要 · Abstract (English)

Large retail outlets offer products that may be domain-specific, and this requires having a model that can understand subtle differences in similar items. Sampling techniques used to train these models are most of the time, computationally expensive or logistically challenging. These models also do not factor in users' previous purchase patterns or behavior, thereby retrieving irrelevant items for them. We present a semantic retrieval model for e-commerce search that embeds queries and products into a shared vector space and leverages a novel taxonomy-based hard-negative sampling(TB-HNS) strategy to mine contextually relevant yet challenging negatives. To further tailor retrievals, we incorporate user-level personalization by modeling each customer's past purchase history and behavior. In offline experiments, our approach outperforms BM25, ANCE and leading neural baselines on Recall@K, while live A/B testing shows substantial uplifts in conversion rate, add-to-cart rate, and average order value. We also demonstrate that our taxonomy-driven negatives reduce training overhead and accelerate convergence, and we share practical lessons from deploying this system at scale.

电商搜索负样本采样个性化推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。