动态调整用户上下文粒度,提升幻想体育推荐效果
Hierarchical Contextual Uplift Bandits for Catalog Personalization
- 分层上下文机制,从全局到个体灵活切换
- 大模型测试中实现0.4%收入提升,用户满意度更高
- 适合需要实时响应的个性化推荐场景
上下文强化学习算法广泛用于个性化推荐,但在幻想体育等动态环境中常因用户行为快速变化和外部因素导致奖励分布剧烈波动而表现不佳,需频繁重训练。为此,我们提出分层上下文增量强化学习框架。该框架通过上下文相似性,动态在系统级粗粒度与用户级细粒度之间调整上下文粒度,促进策略迁移并缓解冷启动问题。同时,融合增量建模思想。在Dream11平台的大规模A/B测试显示,该方法显著提升推荐质量:相比现有生产系统,收入提升0.4%,用户满意度指标改善。2025年5月正式上线后,进一步实现0.5%的收入增长。
原文摘要 · Abstract (English)
Contextual Bandit (CB) algorithms are widely adopted for personalized recommendations but often struggle in dynamic environments typical of fantasy sports, where rapid changes in user behavior and dramatic shifts in reward distributions due to external influences necessitate frequent retraining. To address these challenges, we propose a Hierarchical Contextual Uplift Bandit framework. Our framework dynamically adjusts contextual granularity from broad, system-wide insights to detailed, user-specific contexts, using contextual similarity to facilitate effective policy transfer and mitigate cold-start issues. Additionally, we integrate uplift modeling principles into our approach. Results from large-scale A/B testing on the Dream11 fantasy sports platform show that our method significantly enhances recommendation quality, achieving a 0.4% revenue improvement while also improving user satisfaction metrics compared to the current production system. We subsequently deployed this system to production as the default catalog personalization system in May 2025 and observed a further 0.5% revenue improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。