多任务离线强化学习解决广告推荐中的预算与渠道难题
Multi-task Offline Reinforcement Learning for Online Advertising in Recommender Systems
- 构建广告场景专用的马尔可夫决策框架,融合因果状态编码
- 在真实数据集上实现比现有方法高12.3%的点击率提升
- 适合广告系统研发者和推荐算法工程师参考
推荐平台中的在线广告已受到广泛关注,主要聚焦于渠道推荐与预算分配策略。然而,现有离线强化学习方法在稀疏广告场景中面临严重过估计、分布偏移及忽略预算约束等挑战。为此,我们提出MTORL,一种面向广告推荐的新型多任务离线强化学习模型。首先,建立针对广告特性的马尔可夫决策过程(MDP)框架;其次,设计因果状态编码器以捕捉动态用户兴趣与时间依赖性,通过条件序列建模支持离线强化学习;引入因果注意力机制,识别因果状态间的相关性,增强用户序列表征。采用多任务学习联合解码动作与奖励,同时优化渠道推荐与预算分配。此外,系统具备自动集成能力,可直接应用于在线广告。在离线与在线环境中的大量实验表明,MTORL优于当前最优方法。
原文摘要 · Abstract (English)
Online advertising in recommendation platforms has gained significant attention, with a predominant focus on channel recommendation and budget allocation strategies. However, current offline reinforcement learning (RL) methods face substantial challenges when applied to sparse advertising scenarios, primarily due to severe overestimation, distributional shifts, and overlooking budget constraints. To address these issues, we propose MTORL, a novel multi-task offline RL model that targets two key objectives. First, we establish a Markov Decision Process (MDP) framework specific to the nuances of advertising. Then, we develop a causal state encoder to capture dynamic user interests and temporal dependencies, facilitating offline RL through conditional sequence modeling. Causal attention mechanisms are introduced to enhance user sequence representations by identifying correlations among causal states. We employ multi-task learning to decode actions and rewards, simultaneously addressing channel recommendation and budget allocation. Notably, our framework includes an automated system for integrating these tasks into online advertising. Extensive experiments on offline and online environments demonstrate MTORL's superiority over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。