通过生成式半监督预训练,提升大规模搜索排名效果
Generative Pre-trained Ranking Model with Over-parameterization at Web-Scale (Extended Abstract)
- 用生成式半监督方法预训练排序模型,缓解标注数据不足问题
- 在真实搜索流量中部署,点击率显著提升
- 适合搜索引擎优化与大规模排序系统研发者
学习排序(LTR)广泛用于网页搜索,根据查询词对检索结果进行相关性排序。然而传统LTR模型面临两大挑战:一是缺乏覆盖广泛查询热度的高质量标注查询-网页对数据,难以应对不同热度的查询;二是模型训练不充分,无法生成泛化表示,导致过拟合。为此,我们提出生成式半监督预训练(GS2P)LTR模型。在公开数据集和大规模搜索引擎采集的真实数据集上进行了大量离线实验,并在真实线上搜索系统中部署,显著提升了实际应用效果。
原文摘要 · Abstract (English)
Learning to rank (LTR) is widely employed in web searches to prioritize pertinent webpages from retrieved content based on input queries. However, traditional LTR models encounter two principal obstacles that lead to suboptimal performance: (1) the lack of well-annotated query-webpage pairs with ranking scores covering a diverse range of search query popularities, which hampers their ability to address queries across the popularity spectrum, and (2) inadequately trained models that fail to induce generalized representations for LTR, resulting in overfitting. To address these challenges, we propose a \emph{\uline{G}enerative \uline{S}emi-\uline{S}upervised \uline{P}re-trained} (GS2P) LTR model. We conduct extensive offline experiments on both a publicly available dataset and a real-world dataset collected from a large-scale search engine. Furthermore, we deploy GS2P in a large-scale web search engine with realistic traffic, where we observe significant improvements in the real-world application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。