arXiv:2501.07761cs.LGcs.AI2025-01被引 4

用预测模型加速长期推荐,避免等数周才反馈

Impatient Bandits: Optimizing for the Long-Term Without Delay

  • 构建贝叶斯预测模型融合短期信号与延迟奖励
  • 在播客推荐中实现2个月重复收听率显著提升
  • 适合需要长期效果优化的推荐系统场景

推荐系统日益需提升用户长期满意度。本文将内容探索任务建模为带延迟奖励的多臂老虎机问题。直接等待完整奖励可能需数周,严重拖慢学习速度;而使用短期代理指标又不能准确反映长期目标。为此,我们提出一种基于贝叶斯滤波的预测模型,整合所有历史信息,融合真实奖励与短期替代结果,生成对长期奖励的概率信念。进一步设计了一种新老虎机算法,利用该预测模型快速识别长期成功的优质内容。理论证明了算法的后悔界依赖于‘渐进反馈价值’(Value of Progressive Feedback),这一信息论度量刻画了早期短期指标对长期结果的预测能力。我们在播客推荐任务中验证该方法,目标是推荐能引发用户持续两个月重复收听的内容。实证结果显示,相比仅优化短期代理或仅依赖延迟奖励的方法,本方法在服务数亿用户的推荐系统中通过A/B测试显著优于基线。

原文摘要 · Abstract (English)

Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration task, which we formalize as a bandit problem with delayed rewards. There is an apparent trade-off in choosing the learning signal: waiting for the full reward to become available might take several weeks, slowing the rate of learning, whereas using short-term proxy rewards reflects the actual long-term goal only imperfectly. First, we develop a predictive model of delayed rewards that incorporates all information obtained to date. Rewards as well as shorter-term surrogate outcomes are combined through a Bayesian filter to obtain a probabilistic belief. Second, we devise a bandit algorithm that quickly learns to identify content aligned with long-term success using this new predictive model. We prove a regret bound for our algorithm that depends on the Value of Progressive Feedback, an information-theoretic metric that captures the quality of short-term leading indicators that are observed prior to the long-term reward. We apply our approach to a podcast recommendation problem, where we seek to recommend shows that users engage with repeatedly over two months. We empirically validate that our approach significantly outperforms methods that optimize for short-term proxies or rely solely on delayed rewards, as demonstrated by an A/B test in a recommendation system that serves hundreds of millions of users.

推荐系统多臂老虎机延迟奖励长期优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。