arXiv:2503.07025cs.IRcs.AI2025-03中稿 · AAAI

用用户点击行为推断文档相关性,提升搜索精度

Weak Supervision for Improved Precision in Search Systems

  • 利用点击日志等弱监督信号推断查询-文档相关性
  • 在大规模搜索系统中显著提升排序精度
  • 适合资源有限但需优化搜索效果的团队

现代搜索引擎日益依赖监督学习方法(如Learning to Rank)和海量数据训练深度学习模型,但标注数据的构建既耗时又昂贵。因此,通常使用用户点击和行为日志作为相关性的代理标签。本文提出一种弱监督方法,用于推断查询-文档对的质量,并将其集成到Learning to Rank框架中,显著提升了大规模搜索系统的排序精度。

原文摘要 · Abstract (English)

Labeled datasets are essential for modern search engines, which increasingly rely on supervised learning methods like Learning to Rank and massive amounts of data to power deep learning models. However, creating these datasets is both time-consuming and costly, leading to the common use of user click and activity logs as proxies for relevance. In this paper, we present a weak supervision approach to infer the quality of query-document pairs and apply it within a Learning to Rank framework to enhance the precision of a large-scale search system.

搜索排序弱监督Learning to Rank

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。