arXiv:2508.20798cs.IR2025-08中稿 · CIKM 2025被引 3

考虑用户个性化差异,提升搜索排序模型的公平性。

Addressing Personalized Bias for Unbiased Learning to Rank

  • 引入用户行为偏好与点击倾向的个性化因素,改进排序模型。
  • 在三个数据集上验证,新方法显著降低偏差并提升排序效果。
  • 适合关注搜索公平性与个性化推荐的研究者和工程师。

无偏学习排序(ULTR)旨在从有偏的用户行为日志中学习无偏的排序模型,在网络搜索中具有重要意义。以往研究多关注位置偏差、展示偏差和异常值偏差等常见偏差,但通常假设行为数据来自‘平均’用户,忽略了不同用户在搜索和浏览行为上的差异。本文首次将个性化因素引入ULTR框架,提出用户感知的无偏学习排序问题。通过形式化因果分析,我们证明现有忽略用户差异的方法在用户对查询偏好或文档查看倾向不同时存在偏差。为此,我们提出一种新型用户感知逆倾向评分估计器:针对每个查询建模用户浏览行为分布,并聚合加权的查看概率以确定倾向性。理论证明该估计器在温和假设下无偏且方差更低。大量实验在两个半合成数据集和一个真实数据集上验证了方法的有效性。

原文摘要 · Abstract (English)

Unbiased learning to rank (ULTR), which aims to learn unbiased ranking models from biased user behavior logs, plays an important role in Web search. Previous research on ULTR has studied a variety of biases in users' clicks, such as position bias, presentation bias, and outlier bias. However, existing work often assumes that the behavior logs are collected from an ``average'' user, neglecting the differences between different users in their search and browsing behaviors. In this paper, we introduce personalized factors into the ULTR framework, which we term the user-aware ULTR problem. Through a formal causal analysis of this problem, we demonstrate that existing user-oblivious methods are biased when different users have different preferences over queries and personalized propensities of examining documents. To address such a personalized bias, we propose a novel user-aware inverse-propensity-score estimator for learning-to-rank objectives. Specifically, our approach models the distribution of user browsing behaviors for each query and aggregates user-weighted examination probabilities to determine propensities. We theoretically prove that the user-aware estimator is unbiased under some mild assumptions and shows lower variance compared to the straightforward way of calculating a user-dependent propensity for each impression. Finally, we empirically verify the effectiveness of our user-aware estimator by conducting extensive experiments on two semi-synthetic datasets and a real-world dataset.

学习排序无偏性个性化因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。