用监督学习提升强化学习稳定性,优化直播推荐策略。
Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed
- 融合监督学习与多组状态分解,降低策略更新方差。
- 在离线评估和线上测试中均提升用户留存与使用时长。
- 适合大规模直播推荐系统,尤其关注长期收益的场景。
在短视频与直播混合推荐场景下,直播推荐系统需为每个用户请求决定是否在视频流中插入最多一条直播。为最大化长期用户参与度,需制定最优直播分配策略。不合理的分配策略会显著影响应用使用时长与用户留存,忽略其长期负面影响。近年来,强化学习(RL)被广泛用于捕捉长期用户参与度,但传统算法常面临发散与不稳定问题,限制了其在大规模工业推荐系统中的应用,尤其在此类复杂场景中。为此,本文提出一种新型监督学习增强的多组演员-评论家算法(SL-MGAC)。具体地,引入监督学习增强的演员-评论家框架,结合方差减少技术,多任务奖励学习可抑制评论家学习中的自举误差累积;同时设计多组状态分解模块,分别作用于演员与评论家网络,降低预测方差并提升模型稳定性;还提出新颖奖励函数,防止过度贪婪的直播分配。通过离线策略评估(OPE)与线上A/B测试进行实证评估,结果表明,所提方法在平台级约束下优于基线方法,并在线上推荐场景中展现出更强的稳定性。
原文摘要 · Abstract (English)
In the context of a short video & live stream mixed recommendation scenario, the live stream recommendation system (RS) decides whether to allocate at most one live stream into the video feed for each user request. To maximize long-term user engagement, it is crucial to determine an optimal live stream policy for accurate live stream allocation. The inappropriate live stream allocation policy can significantly affect the duration of the usage app and user retention, which ignores the long-term negative impact of live stream allocation. Recently, reinforcement learning (RL) has been widely applied in recommendation systems to capture long-term user engagement. However, traditional RL algorithms often face divergence and instability problems, which restricts the application and deployment in the large-scale industrial recommendation systems, especially in the aforementioned challenging scenario. To address these challenges, we propose a novel Supervised Learning-enhanced Multi-Group Actor Critic algorithm (SL-MGAC). Specifically, we introduce a supervised learning-enhanced actor-critic framework that incorporates variance reduction techniques, where multi-task reward learning helps restrict bootstrapping error accumulation during critic learning. Additionally, we design a multi-group state decomposition module for both actor and critic networks to reduce prediction variance and improve model stability. We also propose a novel reward function to prevent overly greedy live stream allocation. Empirically, we evaluate the SL-MGAC algorithm using offline policy evaluation (OPE) and online A/B testing. Experimental results demonstrate that the proposed method not only outperforms baseline methods under the platform-level constraints but also exhibits enhanced stability in online recommendation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。