arXiv:2412.06791cs.IRcs.LG2024-12被引 2

用强化学习提升新闻推荐精准度,兼顾个性化与内容更新速度

Enhancing Prediction Models with Reinforcement Learning

  • 融合多臂赌博机与大语言模型的强化学习推荐架构
  • 线上指标显著提升,有效解决冷启动与内容新鲜度问题
  • 适合需要动态调整推荐策略的实时内容平台

我们介绍了在Ringier Axel Springer Polska部署的大规模新闻推荐系统Aureus,通过强化学习技术增强预测模型。该系统整合了多臂赌博机方法和基于大语言模型(LLMs)的深度学习模型,详细阐述了其架构与实现,并强调了将排序预测模型与强化学习结合后在线指标的显著改善。论文进一步探讨了不同模型混合对关键业务绩效指标的影响。该方法有效平衡了个性化推荐与快速适应新闻内容变化的需求,解决了冷启动问题和内容新鲜度挑战。线上评估结果表明,该系统在真实生产环境中具有显著有效性。

原文摘要 · Abstract (English)

We present a large-scale news recommendation system implemented at Ringier Axel Springer Polska, focusing on enhancing prediction models with reinforcement learning techniques. The system, named Aureus, integrates a variety of algorithms, including multi-armed bandit methods and deep learning models based on large language models (LLMs). We detail the architecture and implementation of Aureus, emphasizing the significant improvements in online metrics achieved by combining ranking prediction models with reinforcement learning. The paper further explores the impact of different models mixing on key business performance indicators. Our approach effectively balances the need for personalized recommendations with the ability to adapt to rapidly changing news content, addressing common challenges such as the cold start problem and content freshness. The results of online evaluation demonstrate the effectiveness of the proposed system in a real-world production environment.

推荐系统强化学习新闻推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。