通过分层专家模仿提升推荐系统用户留存率
Stratified Expert Cloning for Retention-Aware Recommendation at Scale
- 按留存行为分层建模专家,动态匹配用户策略
- 线上测试提升0.122%活跃天数,日增超20万用户
- 适合大规模推荐系统长期留存优化场景
用户留存是大规模推荐系统长期成功的关键。现有方法多关注短期互动,忽视用户行为随时间演化的动态性。强化学习虽有望优化长期收益,但面临信用延迟分配和样本效率低的问题。本文提出分层专家模仿(SEC)框架,利用高留存用户的海量交互数据学习稳健策略。SEC包含:1)多层次专家分层以建模多样留存行为;2)自适应专家选择,根据用户状态与留存历史动态匹配合适策略;3)动作熵正则化,提升推荐多样性与策略泛化能力。在快手及快手极速版两大视频平台的离线评估与在线A/B测试中,覆盖数亿用户,结果表明累计活跃天数分别提升0.098%和0.122%,对应每日新增活跃用户超20万。
原文摘要 · Abstract (English)
User retention is critical in large-scale recommender systems, significantly influencing online platforms' long-term success. Existing methods typically focus on short-term engagement, neglecting the evolving dynamics of user behaviors over time. Reinforcement learning (RL) methods, though promising for optimizing long-term rewards, face challenges like delayed credit assignment and sample inefficiency. We introduce Stratified Expert Cloning (SEC), an imitation learning framework that leverages abundant interaction data from high-retention users to learn robust policies. SEC incorporates: 1) multi-level expert stratification to model diverse retention behaviors; 2) adaptive expert selection to dynamically match users with appropriate policies based on their state and retention history; and 3) action entropy regularization to enhance recommendation diversity and policy generalization. Extensive offline evaluations and online A/B tests on major video platforms (Kuaishou and Kuaishou Lite) with hundreds of millions of users validate SEC's effectiveness. Results show substantial improvements, achieving cumulative lifts of 0.098 percent and 0.122 percent in active days on the two platforms respectively, each translating into over 200,000 additional daily active users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。