用离线强化学习优化广告投放,提升用户体验与收益平衡。
Session-Level Dynamic Ad Load Optimization using Offline Robust Reinforcement Learning
- 基于离线DQN框架,缓解动态系统中的混淆偏差问题。
- 相比最佳因果学习基线,离线收益提升超80%,真实数据额外增益5%。
- 适合大规模广告系统、追求稳定收益与用户体验的平台方。
会话级动态广告加载优化旨在用户在线会话中实时个性化调整广告密度与类型,以动态平衡用户体验质量与广告收益。传统基于因果学习的方法在处理混淆偏差和分布偏移等关键技术挑战时表现不佳。本文提出一种基于离线深度Q网络(DQN)的框架,有效缓解动态系统中的混淆偏差,离线性能优于最佳因果学习基线超过80%。为进一步提升对未预见分布偏移的鲁棒性,我们引入新型离线稳健双通道DQN方法,在多个OpenAI-Gym数据集上随扰动增加仍保持更稳定奖励,且在真实广告投放数据上带来额外5%的离线收益。该方法已部署于多个生产系统,上线后在线A/B测试显示,在参与度-广告评分权衡效率上实现两位数提升,显著增强平台服务用户与广告主的能力。
原文摘要 · Abstract (English)
Session-level dynamic ad load optimization aims to personalize the density and types of delivered advertisements in real time during a user's online session by dynamically balancing user experience quality and ad monetization. Traditional causal learning-based approaches struggle with key technical challenges, especially in handling confounding bias and distribution shifts. In this paper, we develop an offline deep Q-network (DQN)-based framework that effectively mitigates confounding bias in dynamic systems and demonstrates more than 80% offline gains compared to the best causal learning-based production baseline. Moreover, to improve the framework's robustness against unanticipated distribution shifts, we further enhance our framework with a novel offline robust dueling DQN approach. This approach achieves more stable rewards on multiple OpenAI-Gym datasets as perturbations increase, and provides an additional 5% offline gains on real-world ad delivery data. Deployed across multiple production systems, our approach has achieved outsized topline gains. Post-launch online A/B tests have shown double-digit improvements in the engagement-ad score trade-off efficiency, significantly enhancing our platform's capability to serve both consumers and advertisers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。