arXiv:2409.09881cs.LGcs.IR2024-09中稿 · the CONSEQUENCES 2…被引 1

提出PRPO方法,在无用户假设下实现推荐系统部署安全。

Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank

  • 通过限制模型与安全基线的差距,避免排名行为过度偏离。
  • 实验显示性能优于现有安全方法,极端情况下仍保持安全。
  • 无需用户行为假设,适合真实场景部署,具有普适安全性。

反事实学习排序(CLTR)在部署时可能存在风险,导致模型性能下降。虽然已有安全机制用于缓解逆倾向评分带来的偏差问题,但现有方法不适用于先进CLTR模型,无法处理信任偏差,且依赖特定用户行为假设。本文提出一种新方法——近端排序策略优化(PRPO),在不依赖用户行为假设的前提下,确保部署安全。PRPO通过抑制模型学习与安全基线差异过大的排序行为,限制其性能退化程度。实验表明,PRPO在性能上优于现有安全逆倾向评分方法,且在最恶劣对抗情形下仍能保证安全。该方法无需假设,是首个实现无条件部署安全并可转化为真实应用鲁棒安全性的方法。

原文摘要 · Abstract (English)

Counterfactual learning to rank (CLTR) can be risky and, in various circumstances, can produce sub-optimal models that hurt performance when deployed. Safe CLTR was introduced to mitigate these risks when using inverse propensity scoring to correct for position bias. However, the existing safety measure for CLTR is not applicable to state-of-the-art CLTR methods, cannot handle trust bias, and relies on specific assumptions about user behavior. We propose a novel approach, proximal ranking policy optimization (PRPO), that provides safety in deployment without assumptions about user behavior. PRPO removes incentives for learning ranking behavior that is too dissimilar to a safe ranking model. Thereby, PRPO imposes a limit on how much learned models can degrade performance metrics, without relying on any specific user assumptions. Our experiments show that PRPO provides higher performance than the existing safe inverse propensity scoring approach. PRPO always maintains safety, even in maximally adversarial situations. By avoiding assumptions, PRPO is the first method with unconditional safety in deployment that translates to robust safety for real-world applications.

排序学习安全优化反事实学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。