arXiv:2508.09636cs.IRcs.AI2025-08

融合表格与非表格数据,用多任务学习提升个性化商品搜索排名。

Personalized Product Search Ranking: A Multi-Task Learning Approach with Tabular and Non-Tabular Data

  • 通过多任务学习融合表格与非表格数据,使用TinyBERT生成语义嵌入。
  • 在真实数据集上,相比基线模型提升点击率与相关性指标。
  • 适合电商推荐系统研发者,尤其关注个性化排序与标注效率优化。

本文提出一种新型多任务学习框架,用于优化个性化商品搜索排名。该方法创新性地整合表格数据与非表格数据,采用预训练的TinyBERT模型生成语义嵌入,并引入一种新采样策略以捕捉多样化的用户行为。实验对比了XGBoost、TabNet、FT-Transformer、DCN-V2和MMoE等基线模型,重点评估其处理混合数据类型及优化个性化排序的能力。此外,我们设计了一种基于点击率、点击位置和语义相似度的可扩展相关性标注机制,替代传统人工标注。实验结果表明,在多任务学习范式下结合非表格数据与先进嵌入技术能显著提升模型性能。消融实验进一步验证了相关性标签、Fine-tune TinyBERT层以及查询-商品嵌入交互的有效性。这些结果证明了该方法在提升个性化商品搜索排名方面的有效性。

原文摘要 · Abstract (English)

In this paper, we present a novel model architecture for optimizing personalized product search ranking using a multi-task learning (MTL) framework. Our approach uniquely integrates tabular and non-tabular data, leveraging a pre-trained TinyBERT model for semantic embeddings and a novel sampling technique to capture diverse customer behaviors. We evaluate our model against several baselines, including XGBoost, TabNet, FT-Transformer, DCN-V2, and MMoE, focusing on their ability to handle mixed data types and optimize personalized ranking. Additionally, we propose a scalable relevance labeling mechanism based on click-through rates, click positions, and semantic similarity, offering an alternative to traditional human-annotated labels. Experimental results show that combining non-tabular data with advanced embedding techniques in multi-task learning paradigm significantly enhances model performance. Ablation studies further underscore the benefits of incorporating relevance labels, fine-tuning TinyBERT layers, and TinyBERT query-product embedding interactions. These results demonstrate the effectiveness of our approach in achieving improved personalized product search ranking.

个性化推荐多任务学习商品搜索嵌入表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。