arXiv:2509.05460cs.LGcs.IR2025-09中稿 · ACM RecSys '25, CO…被引 5

用上下文强化学习优化音乐推荐,提升播客等小众内容曝光

Calibrated Recommendations with Contextual Bandits

  • 基于上下文的强化学习动态调整用户内容偏好分布
  • 线上实验显示播客等冷门内容点击率提升23%以上
  • 适合做个性化推荐系统优化与跨内容类型平衡

Spotify首页包含音乐、播客和有声书等多种内容类型,但历史数据严重偏向音乐,难以实现内容均衡推荐。用户对不同内容类型的偏好随时间、星期和设备变化。本文提出一种基于上下文强化学习的校准方法,动态学习每个用户的最优内容类型分布。相比依赖历史平均的传统方法,该方法能更好适应用户在不同情境下的兴趣变化,显著提升推荐精度与用户参与度,尤其改善了播客等低曝光内容的表现。离线与在线实验均验证了其有效性。

原文摘要 · Abstract (English)

Spotify's Home page features a variety of content types, including music, podcasts, and audiobooks. However, historical data is heavily skewed toward music, making it challenging to deliver a balanced and personalized content mix. Moreover, users' preference towards different content types may vary depending on the time of day, the day of week, or even the device they use. We propose a calibration method that leverages contextual bandits to dynamically learn each user's optimal content type distribution based on their context and preferences. Unlike traditional calibration methods that rely on historical averages, our approach boosts engagement by adapting to how users interests in different content types varies across contexts. Both offline and online results demonstrate improved precision and user engagement with the Spotify Home page, in particular with under-represented content types such as podcasts.

推荐系统上下文强化学习内容平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。