arXiv:2607.14192cs.LGcs.IR2026-07被引 1

不依赖特定模型,自动挖掘用户行为预测长期留存。

Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning

论文配图:Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning
图 1 · 摘自论文原文
  • 从用户早期行为中筛选可预测留存的可观测动作
  • 设计多个无需模型改造的下游奖励信号,提升长期留存率
  • 已在Pinterest多场景上线,效果稳定且易部署

近年来,推荐系统优化目标已从短期行为转向长期用户参与和留存。但直接优化留存困难,因回报信号稀疏、延迟且难以归因于早期推荐。此前工作采用序列建模和强化学习,但需定制奖励设计、计算开销大,通用性差。本文提出一种统一的、模型无关的下游奖励学习框架,用于大规模推荐系统中的长期用户价值优化。首先,我们定义下游奖励学习问题,并构建离线筛选框架,识别出早期可观测且能预测未来留存的会话级行为。随后,基于多源用户行为模式,提出若干模型无关的下游奖励信号。我们还讨论了将这些奖励信号工程化并集成到排序模型中的挑战与实践。在线A/B实验显示,各项参与度和留存指标均持续提升,该框架已在Pinterest的Homefeed、Related Pins、Search及Notifications等多个场景部署。

原文摘要 · Abstract (English)

As recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention. However, directly optimizing retention is difficult because return signals are sparse, delayed, and only partially attributable to earlier recommendations. Prior work has addressed this challenge with sequential modeling and reinforcement learning, but these approaches typically require task specific reward engineering, substantial computational overhead, and surface specific implementations that are difficult to generalize. In this paper, we present a unified, model-agnostic downstream reward framework for optimizing long-term user value in large-scale recommendation systems. First, we formulate the downstream reward learning problem and develop an offline screening framework to identify session level behaviors that are both observable early and predictive of future retention. We then propose several model-agnostic downstream rewards signals derived from observed user action patterns across multiple sources. We further discuss the engineering effort to productionize the proposed rewards derivations and challenges we faced when adding them to our ranking models. Online A/B experiments demonstrate consistent improvements in engagement and retention-related metrics, and the framework has been deployed across multiple Pinterest surfaces, including Homefeed, Related Pins, Search, and Notifications.

推荐系统长期留存奖励学习模型无关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。