arXiv:2409.12043cs.IRcs.LG2024-09中稿 · the CONSEQUENCES '…

实证发现双塔模型在真实数据中不受日志策略混淆影响

Understanding the Effects of the Baidu-ULTR Logging Policy on Two-Tower Models

  • 在百度真实数据集上验证混淆问题存在性
  • 双塔模型性能未受日志策略混淆显著影响
  • 专家标注与用户点击行为存在潜在偏差

尽管双塔模型在无偏学习排序(ULTR)任务中广受欢迎,但近期研究指出其存在重大局限性,可能导致在工业应用中失效:日志策略混淆问题。已有多种解决方案被提出,但评估多基于半合成仿真实验。本文通过分析最大规模的真实数据集Baidu-ULTR,填补了理论与实践的鸿沟。主要贡献包括:1)证实了在Baidu-ULTR数据集中存在混淆问题的条件;2)发现该混淆问题对双塔模型性能无显著影响;3)揭示了专家标注(ULTR中的黄金标准)与用户点击行为之间可能存在不匹配。

原文摘要 · Abstract (English)

Despite the popularity of the two-tower model for unbiased learning to rank (ULTR) tasks, recent work suggests that it suffers from a major limitation that could lead to its collapse in industry applications: the problem of logging policy confounding. Several potential solutions have even been proposed; however, the evaluation of these methods was mostly conducted using semi-synthetic simulation experiments. This paper bridges the gap between theory and practice by investigating the confounding problem on the largest real-world dataset, Baidu-ULTR. Our main contributions are threefold: 1) we show that the conditions for the confounding problem are given on Baidu-ULTR, 2) the confounding problem bears no significant effect on the two-tower model, and 3) we point to a potential mismatch between expert annotations, the golden standard in ULTR, and user click behavior.

双塔模型无偏排序日志混淆真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。