arXiv:2507.16253cs.IR2025-07

通过强化学习提升用户与作者的长期互动价值,助力短视频平台生态繁荣。

Reinforce Lifelong Interaction Value of User-Author Pairs for Large-Scale Recommendation Systems

  • 构建稀疏跨请求交互马尔可夫决策过程,解决长间隔与大规模交互难题
  • 多任务批评学习捕捉从点击到打赏的渐进式互动信号,提升稀疏标签学习效果
  • 适合关注社交平台长期生态、推荐系统优化的研究者与工程师

推荐系统帮助用户发现感兴趣内容并连接作者与目标受众。现有研究多聚焦于精准预测用户即时反馈(如点击率)或提升长期参与度,却忽视了对作者的影响及用户-作者对的终生互动价值(LIV),这对短视频平台的社区繁荣至关重要。强化学习(RL)擅长优化长期收益,已被广泛应用于推荐系统。本文提出基于用户-作者对每次互动的强化学习方法——RLIV-UA,以增强其终生互动价值。针对用户-作者互动间隔长、空间规模大的问题,设计了稀疏跨请求交互马尔可夫决策过程(SCRI-MDP),并引入邻近状态近似(ASA)构造训练样本。同时,提出多任务批评学习(MTCL),利用密集互动信号(如点击→关注→打赏)补偿稀疏标签的学习。此外,设计辅助监督学习任务以加速模型收敛。离线实验与在线A/B测试均表明,RLIV-UA在提升用户满意度和平台收益方面优于对比方法。

原文摘要 · Abstract (English)

Recommendation systems (RS) help users find interested content and connect authors with their target audience. Most research in RS tends to focus either on predicting users' immediate feedback (like click-through rate) accurately or improving users' long-term engagement. However, they ignore the influence for authors and the lifelong interaction value (LIV) of user-author pairs, which is particularly crucial for improving the prosperity of social community in short-video platforms. Currently, reinforcement learning (RL) can optimize long-term benefits and has been widely applied in RS. In this paper, we introduce RL to Reinforce Lifelong Interaction Value of User-Author pairs (RLIV-UA) based on each interaction of UA pairs. To address the long intervals between UA interactions and the large scale of the UA space, we propose a novel Sparse Cross-Request Interaction Markov Decision Process (SCRI-MDP) and introduce an Adjacent State Approximation (ASA) method to construct RL training samples. Additionally, we introduce Multi-Task Critic Learning (MTCL) to capture the progressive nature of UA interactions (click -> follow -> gift), where denser interaction signals are leveraged to compensate for the learning of sparse labels. Finally, an auxiliary supervised learning task is designed to enhance the convergence of the RLIV-UA model. In offline experiments and online A/B tests, the RLIV-UA model achieves both higher user satisfaction and higher platform profits than compared methods.

推荐系统强化学习长时互动短视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。